mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
merge: absorb upstream drift b9f7597f..8d13373b into v7next (F6 rolling sync)
Upstream = semantic truth, campaign = structural truth. Every upstream
semantic delta lands in its campaign owner leaf; upstream duplicate
extractions do not survive as twins:
- acceptance_dialogue.py -> folded into loop_acceptance{,_review}.py
(A-material paid identity, free replay, identical-refusal terminal,
dialogue history, inconclusive-dialogue reducer semantics)
- delivery_protocol.py -> folded into loop_delivery.py (hold-control
literals, RecursionError-degraded and trailing-object protocol parsers
over the shared strip_protocol_fence normalization)
- chat_delivery_events.py -> folded into events_chat_delivery.py
(unified _delivery_chat_id incl. chat-0 media, send_links/send_quiz,
registry merged via **_CDE)
- events review-wave handler retired for telemetry_events.py registry
- python_interpreter.py -> process_interpreters.py (upstream as-is);
registry/tool_resolution/shell retargeted; interpreter_attestation
scope in registry_core; node post-gates predispatch
- R5 typed process facts: process_facts.py channel + loop-side merge
into the typed result_meta; describe_returncode SSOT retargeted
- deadline_utils.deadline_expired public name adopted (rename-class);
test pins moved off the private spelling
- protected surfaces (BIBLE.md, safety.py, CHECKLISTS.md, registry.py,
gateway/contracts.py incl. endpoint_index extraction) landed as-is
- web wave, VERSION 6.113.5, package data landed as-is
- size_ratchet_manifest regenerated via scripts/regenerate_size_ratchet.py
(band rationales recorded for the F6-grown leaves); tools/core.py
link/quiz/escalate spans moved to core_artifacts.py to stay under the
giant gate; _run_shell and registry dispatch shaved under the
function gate via shell_process/registry_core helpers
This commit is contained in:
commit
2085019152
351 changed files with 32079 additions and 2896 deletions
30
.github/PULL_REQUEST_TEMPLATE.md
vendored
30
.github/PULL_REQUEST_TEMPLATE.md
vendored
|
|
@ -45,9 +45,10 @@ Evidence:
|
|||
|
||||
## Governance and documentation
|
||||
|
||||
- [ ] I read `CONTRIBUTING.md`; for a substantive change I also read
|
||||
`BIBLE.md`, `docs/ARCHITECTURE.md`, `docs/DEVELOPMENT.md`, and
|
||||
`docs/CHECKLISTS.md` in full.
|
||||
- [ ] I read `CONTRIBUTING.md` and `docs/CHECKLISTS.md` in full; for a
|
||||
substantive change I mapped `BIBLE.md`, `docs/ARCHITECTURE.md`,
|
||||
`docs/DEVELOPMENT.md`, and `docs/DESIGN.md` by their headings and read
|
||||
every section relevant to this change in full.
|
||||
- [ ] I updated tests and documentation where behavior or architecture changed.
|
||||
- [ ] I did not include secrets, local settings, runtime state, logs, caches, or
|
||||
generated build/review artifacts in the commit.
|
||||
|
|
@ -75,6 +76,29 @@ available, use NOT_RUN and say why.
|
|||
- Full review output or artifact link:
|
||||
- If not run, reason:
|
||||
|
||||
Scope checklist coverage (from the reviewer's JSON; one row per item, extra
|
||||
rows for additional FAIL findings on the same item):
|
||||
|
||||
| Item | Verdict | Evidence |
|
||||
| --- | --- | --- |
|
||||
| intent_alignment | | |
|
||||
| forgotten_touchpoints | | |
|
||||
| cross_surface_consistency | | |
|
||||
| regression_surface | | |
|
||||
| prompt_doc_sync | | |
|
||||
| architecture_fit | | |
|
||||
| cross_module_bugs | | |
|
||||
| implicit_contracts | | |
|
||||
|
||||
<details>
|
||||
<summary>Reviewer checklist JSON (validate with scripts/validate_scope_receipt.py)</summary>
|
||||
|
||||
```json
|
||||
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Final checklist
|
||||
|
||||
- [ ] The PR base branch is `ouroboros`.
|
||||
|
|
|
|||
3
BIBLE.md
3
BIBLE.md
|
|
@ -597,6 +597,9 @@ its code in a single session.
|
|||
file-size budgets)
|
||||
- [docs/CHECKLISTS.md](docs/CHECKLISTS.md) — review checklists and
|
||||
plan-review triggers (SSOT referenced by the immune system, P3)
|
||||
- [docs/DESIGN.md](docs/DESIGN.md) — visual and interaction
|
||||
semantics of the UI (engineering rules live in
|
||||
docs/DEVELOPMENT.md § Design System)
|
||||
- [ouroboros/config.py](ouroboros/config.py) — runtime defaults
|
||||
- `memory/knowledge/patterns.md` (under the runtime data root;
|
||||
created on first write) — Pattern Register projection
|
||||
|
|
|
|||
|
|
@ -19,8 +19,12 @@ this flow.
|
|||
|
||||
## 1. Read the Project Before Editing
|
||||
|
||||
For a substantive change, read these files **in full** before designing or
|
||||
editing:
|
||||
For a substantive change, ground yourself in the project documents before
|
||||
designing or editing. Read [`docs/CHECKLISTS.md`](docs/CHECKLISTS.md) — the
|
||||
review checklist single source of truth — **in full**: your change will be
|
||||
judged against it. For the other four, build a navigation map from their
|
||||
headings first, then read every section relevant to your change **in full**
|
||||
(skimming a relevant section does not count):
|
||||
|
||||
- [`BIBLE.md`](BIBLE.md) — constitutional principles and design priorities.
|
||||
- [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) — structure, data flows, and
|
||||
|
|
@ -28,16 +32,18 @@ editing:
|
|||
- [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md) — engineering, testing, and
|
||||
review conventions.
|
||||
- [`docs/DESIGN.md`](docs/DESIGN.md) — visual and interaction semantics.
|
||||
- [`docs/CHECKLISTS.md`](docs/CHECKLISTS.md) — the review checklist single
|
||||
source of truth.
|
||||
|
||||
When in doubt whether a section is relevant, read it. Reading everything in
|
||||
full remains the strongest preparation for a large or cross-cutting change.
|
||||
|
||||
Reuse the modules, contracts, and authorities those documents name. Do not
|
||||
invent a parallel mechanism when the repository already has one. A useful
|
||||
first instruction for a coding agent is:
|
||||
|
||||
> Read CONTRIBUTING.md, BIBLE.md, docs/ARCHITECTURE.md,
|
||||
> docs/DEVELOPMENT.md, docs/DESIGN.md, and docs/CHECKLISTS.md in full before
|
||||
> editing. Follow their current architecture and keep the requested change focused.
|
||||
> Read CONTRIBUTING.md and docs/CHECKLISTS.md in full. Map BIBLE.md,
|
||||
> docs/ARCHITECTURE.md, docs/DEVELOPMENT.md, and docs/DESIGN.md by their
|
||||
> headings and read every section relevant to the requested change in full.
|
||||
> Follow their current architecture and keep the requested change focused.
|
||||
|
||||
These documents may themselves be improved, but constitutional changes must
|
||||
follow the semantic change process in `BIBLE.md`, and behavior, tests, and
|
||||
|
|
@ -124,6 +130,16 @@ Before opening a substantive PR, the authoring agent must hand the final
|
|||
committed diff to a **separate agent context**. Use a subagent, new task, or
|
||||
fresh agent session. Reviewing in the authoring conversation does not count.
|
||||
|
||||
The main review path is an **agentic checklist review**: the reviewer reads
|
||||
the repository with its own tools and covers the "Intent / Scope Review
|
||||
Checklist" from [`docs/CHECKLISTS.md`](docs/CHECKLISTS.md) — every one of its
|
||||
eight items (`intent_alignment`, `forgotten_touchpoints`,
|
||||
`cross_surface_consistency`, `regression_surface`, `prompt_doc_sync`,
|
||||
`architecture_fit`, `cross_module_bugs`, `implicit_contracts`) — following
|
||||
that checklist's output contract: one JSON array covering all eight items,
|
||||
PASS rows mandatory and justified with a concrete artifact, FAIL rows with
|
||||
severity (a critical FAIL names an exact file/symbol).
|
||||
|
||||
Give the reviewer the issue or goal, non-goals, exact base and head SHAs, and
|
||||
repository access. The reviewer must not edit the candidate. Use this compact
|
||||
instruction:
|
||||
|
|
@ -131,16 +147,33 @@ instruction:
|
|||
```text
|
||||
Review the final pull request diff from <base SHA> to <head SHA>. Do not edit.
|
||||
|
||||
Read CONTRIBUTING.md. For a substantive change, read BIBLE.md,
|
||||
docs/ARCHITECTURE.md, docs/DEVELOPMENT.md, docs/DESIGN.md, and
|
||||
docs/CHECKLISTS.md in full.
|
||||
Read CONTRIBUTING.md and docs/CHECKLISTS.md in full. Map BIBLE.md,
|
||||
docs/ARCHITECTURE.md, docs/DEVELOPMENT.md, and docs/DESIGN.md by their
|
||||
headings and read every section relevant to this change in full.
|
||||
Inspect the complete diff, touched files, relevant callers, tests, and docs.
|
||||
|
||||
Return findings first with severity and exact file/line references, then checks
|
||||
Cover the Intent / Scope Review Checklist from docs/CHECKLISTS.md exactly:
|
||||
output a JSON array of objects with the keys "item", "verdict" (PASS/FAIL),
|
||||
"severity" (critical/advisory), and "reason", covering all eight checklist
|
||||
items per its output contract — PASS rows are mandatory and justified with a
|
||||
concrete artifact; a critical FAIL names an exact file/symbol. Then report
|
||||
any further findings with severity and file/line references, checks
|
||||
performed, coverage limitations, and one verdict: PASS, NEEDS_CHANGES, or
|
||||
INCOMPLETE. Do not report PASS when required material was unavailable.
|
||||
```
|
||||
|
||||
Paste the reviewer's checklist JSON into the PR's review-evidence block and
|
||||
fill the checklist table from it. Validate the JSON locally before opening
|
||||
the PR — the schema validator reuses the runtime contract and accepts the
|
||||
array bare, inside a fenced `json` code block, or embedded in the
|
||||
reviewer's prose. It
|
||||
validates SHAPE, not truth: a passing receipt is well-formed coverage, not a
|
||||
review verdict.
|
||||
|
||||
```bash
|
||||
python scripts/validate_scope_receipt.py path/to/receipt.json
|
||||
```
|
||||
|
||||
Reproduce material findings when possible. Fix confirmed problems, and briefly
|
||||
record why any finding was rejected or deferred. Any code change, rebase, or
|
||||
conflict resolution makes the old review stale; review the new final range.
|
||||
|
|
@ -150,12 +183,26 @@ loop.
|
|||
If no separate agent context is available, do not substitute same-context
|
||||
self-review. Mark the review `NOT_RUN` and explain why in the PR.
|
||||
|
||||
### Optional project-native review command
|
||||
### Maintainer-grade project-native review command
|
||||
|
||||
Ouroboros can produce the same evidence in a structured SHA-bound packet. Its
|
||||
Ouroboros can produce review evidence in a structured SHA-bound packet. Its
|
||||
contributor mode uses the reviewer slots actually configured on the machine:
|
||||
`api_chat`, `agent_session`, or a mixture.
|
||||
|
||||
Treat this command as **maintainer / large-window tooling**, not the default
|
||||
contributor path. The scope reviewer's required-artifact pack (protected
|
||||
runtime paths, prompts, contracts, canonical docs, the review stack) is
|
||||
required regardless of how small the diff is, and on a default install it
|
||||
can exceed the configured scope slot's context window even after every
|
||||
degradation step — the run then fails closed with `SCOPE_REVIEW_BLOCKED`
|
||||
and still preserves the evidence packet (marked incomplete). The documented
|
||||
routes past that pack budget are: configure the scope row as an
|
||||
`agent_session` reviewer (a different delivery class — it reads the
|
||||
repository with its own tools instead of being handed one assembled pack,
|
||||
and needs its own confirmed 200K+ window), or configure an API scope slot
|
||||
whose confirmed context window fits the pack. The agentic checklist review
|
||||
above needs neither.
|
||||
|
||||
Configured API slots need their provider credentials and a positive finite
|
||||
`TOTAL_BUDGET`. Agent-session slots need their configured agent route and
|
||||
account to be available. The wrapper checks route-specific readiness where it
|
||||
|
|
@ -200,6 +247,7 @@ Complete the PR template with:
|
|||
- reviewed base SHA and head SHA;
|
||||
- verdict, findings, and their disposition;
|
||||
- checks, coverage limitations, and full output or artifact link;
|
||||
- the scope-checklist coverage table and the reviewer's checklist JSON;
|
||||
- `NOT_RUN` plus a reason for unavailable verification or review.
|
||||
|
||||
Review output is public evidence. Inspect attachments for credentials, private
|
||||
|
|
@ -213,7 +261,8 @@ commit, push, merge, release, or publication.
|
|||
|
||||
- [ ] The PR targets `ouroboros` and is current with its recorded base.
|
||||
- [ ] The PR has one coherent purpose and explicit non-goals.
|
||||
- [ ] The required project documents were read in full.
|
||||
- [ ] `docs/CHECKLISTS.md` was read in full and every relevant section of the
|
||||
other project documents was read in full.
|
||||
- [ ] Relevant tests and UI evidence are recorded honestly.
|
||||
- [ ] No release-version carrier was changed.
|
||||
- [ ] No secret, runtime state, generated run, or build artifact is in the diff.
|
||||
|
|
|
|||
20
README.md
20
README.md
|
|
@ -12,7 +12,7 @@
|
|||
[](https://ouroboros-agent.ai/install/#linux)
|
||||
[][download-windows-x64]
|
||||
[](https://github.com/razzant/OuroborosHub)
|
||||
[](VERSION)
|
||||
[](VERSION)
|
||||
|
||||
Ouroboros is an open-source, general-purpose AI agent whose identity, durable memory, and history continue across tasks and restarts. It works on external projects, coordinates a live swarm of specialist agents, and can rewrite the implementation it runs on, including its code, architecture, prompts, tools, and dependencies. Reflection can also change how it understands itself without severing that continuity.
|
||||
|
||||
|
|
@ -64,13 +64,13 @@ The desktop packages already contain an optional CLI installer. On macOS, after
|
|||
|
||||
</details>
|
||||
|
||||
[download-macos-arm64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4.dmg
|
||||
[download-windows-x64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-windows-x64.zip
|
||||
[download-linux-deb-amd64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros_6.113.4_amd64.deb
|
||||
[download-linux-rpm-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros-6.113.4-1.x86_64.rpm
|
||||
[download-linux-rpm-red80-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros-6.113.4-1.red80.x86_64.rpm
|
||||
[download-linux-appimage-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-linux-x86_64.AppImage
|
||||
[download-linux-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-linux-x86_64.tar.gz
|
||||
[download-macos-arm64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5.dmg
|
||||
[download-windows-x64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-windows-x64.zip
|
||||
[download-linux-deb-amd64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros_6.113.5_amd64.deb
|
||||
[download-linux-rpm-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros-6.113.5-1.x86_64.rpm
|
||||
[download-linux-rpm-red80-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros-6.113.5-1.red80.x86_64.rpm
|
||||
[download-linux-appimage-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-linux-x86_64.AppImage
|
||||
[download-linux-x86_64]: https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-linux-x86_64.tar.gz
|
||||
|
||||
Ouroboros bundles [Claudexor](https://github.com/razzant/claudexor) as its local execution layer for delegated coding and hosted-agent review. Ouroboros owns the task, memory, review, and final integration, while Claudexor runs the selected connected coding harness and returns durable execution evidence. [Explore Claudexor](https://claudexor.ai/).
|
||||
|
||||
|
|
@ -308,7 +308,7 @@ ouroboros run --start \
|
|||
|
||||
Use `--jsonl` for a machine-readable event stream and `--detach` when the caller will follow the task with `ouroboros tasks watch <task_id>` or inspect it with `ouroboros tasks show <task_id>`. External workspace runs keep Ouroboros's own repository and governance context separate, then export changes as reviewable patch artifacts.
|
||||
|
||||
To change Ouroboros itself, follow [CONTRIBUTING.md](CONTRIBUTING.md) and read [BIBLE.md](BIBLE.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md), [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md), and [docs/CHECKLISTS.md](docs/CHECKLISTS.md) in full before editing.
|
||||
To change Ouroboros itself, follow [CONTRIBUTING.md](CONTRIBUTING.md): read [docs/CHECKLISTS.md](docs/CHECKLISTS.md) in full, and map [BIBLE.md](BIBLE.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md), [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md), and [docs/DESIGN.md](docs/DESIGN.md) by their headings, reading every section relevant to your change in full before editing.
|
||||
|
||||
#### Configuration
|
||||
|
||||
|
|
@ -449,6 +449,7 @@ and the reason.
|
|||
|
||||
| Version | Date | Description |
|
||||
|---------|------|-------------|
|
||||
| 6.113.5 | 2026-08-31 | **fix: browser cleanup after a timed-out tool is generation-safe (integrates community PR #429 by @mikemikimike, closes #409).** A stateful-tool timeout now retires the whole browser generation: the shared state slot is replaced with a fresh object, the abandoned worker keeps writing only into its retired one, and the close is queued on the retiring executor so it always runs on the owning worker thread — including the already-settled race — with the cognitive lease closing after that cleanup. A late infrastructure-error retry that observes a replaced generation closes only its own retired session, so it can no longer cross-thread-kill the next command's browser. Closes the follow-up findings of the #440 post-merge audit; the never-settling-worker session leak stays a disclosed residual of in-process Playwright. |
|
||||
| 6.113.4 | 2026-08-30 | **fix: release verification follows the family-mounted sign-in card.** The browser acceptance tests now re-resolve the active per-family login host after every account-list rebuild, matching the placement introduced by the roster UI while preserving the full cancellation, reconciliation, stale-poll, and terminal-race assertions. No lifecycle check is weakened and the managed Claudexor runtime remains pinned to 3.9.2. |
|
||||
| 6.113.3 | 2026-08-30 | **fix: Claude subscription quotas recover automatically without a user action.** The managed Claudexor runtime advances to 3.9.2. Before an expired or near-expiry access token is used for quota reading, Claude Code now refreshes its own credential through a prompt-free vendor lifecycle: no MCP request, model inference, manual login, refresh-token custody, or direct store write is added to Claudexor. A still-valid token remains available if the proactive wake fails, expired credentials stay typed as unknown rather than falsely logged out, and the exact helper process is gracefully reaped before quota polling continues. |
|
||||
| 6.113.2 | 2026-08-30 | **fix: account truth, operator control, route presentation, and network recovery converge with the Claudexor 3.9.1 runtime.** Claude profiles with refreshable expired or unknown-expiry credentials no longer become falsely logged out when the optional usage endpoint cannot provide quota; typed quota absence stays neutral, percentages remain evidence-backed, and the Accounts row keeps login actions separate from quota refresh. The compact account layout now follows the existing 980px narrow-shell boundary. OpenRouter requests share the canonical Ouroboros app identity, settings writes retain explicit owner and shape authority, route chips name the actual selected path, and task stop controls plus owned-daemon startup report the correct menu and Node probe facts. Proven pre-dispatch network failures now wait and redial without duplicate billing, while the web UI reconnects in place and bounded network operations survive VPN and sleep transitions. The managed Claudexor pin advances to 3.9.1 so these fixes reach every packaged desktop after update and daemon restart. |
|
||||
|
|
@ -458,7 +459,6 @@ and the reason.
|
|||
| 6.110.1 | 2026-08-25 | **fix: ship truthful cross-surface state, review custody, and cross-platform UI and cognition convergence.** OuroborosHub skill cards and publication now follow identity-first sync with durable receipts and an explicit adopt transaction (PR #313). Chat and Project activity reconcile stale Working states from closed lineage and durable task truth, preserve Project-thread routing, and keep the Node gate authoritative (PR #314). Review outputs retain durable custody, swarm work stays visible, the advisory lane is restored, and Projects receive a stable entry point (PR #316). Background Consciousness now derives decoded text and its verification digest from one raw snapshot, so Windows newline normalization or a concurrent rewrite cannot bind observed text to different bytes (PR #321). Widget frames converge their measured geometry without nested overflow ownership fighting the host across WebKit-style engines (PR #320). |
|
||||
| 6.110.0 | 2026-08-22 | **feat: preserve work-order authority across oversized delegation and recovery.** Complete external work orders remain byte-complete within the host serializer budget; when an order needs bounded continuation, Ouroboros asks the same actor for an exact readable canonical range and keeps incomplete coverage as typed `cannot_verify` evidence. Partial input can no longer authorize PASS, a destructive rewrite, or replacement of the full contract, while valid complete work continues through the existing flow. |
|
||||
| 6.109.0 | 2026-08-21 | **feat: live task cost and ready-on-open agent accounts.** Running root-task heartbeats now project the existing physical-attempt ledger into one non-final subtree total, so compact Chat and Activity cards advance live without a second timer, endpoint, or client-side sum while preserving reserved, unresolved, and unmetered disclosure (PR #288). Opening Agents now wakes only an already-provisioned stale Claudexor home through the existing owner-action endpoint after a side-effect-free status read; background polling, first-time installs, foreign homes, and repair states remain untouched (PR #289). Fail-closed staged-binary review fixtures now inject exact Git tree-read errors instead of assuming loose object storage, removing the macOS stable-CI race without changing production behavior. |
|
||||
| 6.107.0 | 2026-08-21 | **feat: secret-safe skill publishing and portable provider tool contracts.** OuroborosHub publishing now works from an immutable reviewed-byte snapshot, scans the exact candidate through pinned Betterleaks before any GitHub effect, keeps raw findings out of model and durable-result surfaces, guides repair and fresh review through the existing managed-task flow, and treats only a validated pull-request receipt as publication success. The Skills UI exposes the selected-skill preflight and publishing journey without turning its passive card projection into a second readiness authority. Provider compatibility now validates the complete shipped built-in tool registry against portable JSON Schema rules and trusted live canaries across OpenRouter, direct OpenAI, direct Anthropic, GigaChat, and configured optional providers; confirmed contract failures block release preflight while pull-request CI remains secretless. |
|
||||
Older releases are preserved in Git tags and GitHub releases. Older 6.x rows (including 6.108.1, 6.106.0, 6.101.1, 6.97.2, 6.105.0, 6.97.1, 6.97.0, 6.96.1, 6.96.0, 6.95.0, 6.94.0, 6.93.0, 6.92.1, 6.92.0, 6.91.1, 6.90.3, 6.91.0, 6.90.2, 6.90.0, 6.87.5, 6.87.4, 6.87.3, 6.87.2, 6.84.0, 6.87.1, 6.83.0, 6.86.1, 6.81.1, 6.76.0, 6.75.0, 6.74.5, 6.74.4, 6.74.1, 6.74.0, 6.73.2, 6.73.1, 6.73.0, 6.72.0, 6.71.2, 6.71.1, 6.71.0, 6.70.0, 6.69.0, 6.68.0, 6.67.0, 6.66.0, 6.65.4, 6.65.3, 6.65.2, 6.65.1, 6.65.0, 6.64.3, 6.64.2, 6.64.1, 6.64.0, 6.63.0, 6.62.0, 6.61.4, 6.61.3, 6.61.1, 6.61.0, 6.60.0, 6.59.0, 6.58.0, 6.57.0, 6.56.0, 6.55.0, 6.54.4, 6.54.2, 6.54.1, 6.54.0, 6.53.4, 6.53.0, 6.51.0), the 5.2.0 through 5.33.0-rc.6 rows, and former `4.0.0` rows are rolled off to respect the P9 changelog cap; their full bodies remain at their git tags.
|
||||
|
||||
---
|
||||
|
|
|
|||
2
VERSION
2
VERSION
|
|
@ -1 +1 @@
|
|||
6.113.4
|
||||
6.113.5
|
||||
|
|
|
|||
|
|
@ -212,6 +212,11 @@ def _detect_version(events: list[dict[str, Any]]) -> str | None:
|
|||
def _final_answer(agent_dir: Path) -> str:
|
||||
chat = _read_jsonl_chain(agent_dir / "ouroboros-data", "chat.jsonl", "chat")
|
||||
for row in reversed(chat):
|
||||
if row.get("type"):
|
||||
# Typed structural rows (quiz/links/document/photo/video,
|
||||
# proactive_message, project lifecycle) are deliveries, never the
|
||||
# task's final answer.
|
||||
continue
|
||||
if row.get("direction") == "out" and isinstance(row.get("text"), str):
|
||||
return row["text"]
|
||||
try:
|
||||
|
|
|
|||
|
|
@ -42,7 +42,7 @@ from devtools.benchmarks.common.run_roots import (
|
|||
safe_join_under,
|
||||
)
|
||||
from devtools.benchmarks.terminal_bench.run_harbor_smoke import AGENT_IMPORT
|
||||
from ouroboros.config import SETTINGS_DEFAULTS
|
||||
from ouroboros.config import EFFORT_SCALE, SETTINGS_DEFAULTS
|
||||
|
||||
|
||||
DEFAULT_DATASET = "terminal-bench/terminal-bench-2-1"
|
||||
|
|
@ -900,7 +900,7 @@ def _build_arg_parser() -> argparse.ArgumentParser:
|
|||
parser.add_argument(
|
||||
"--review-effort",
|
||||
default="low",
|
||||
choices=["none", "low", "medium", "high"],
|
||||
choices=list(EFFORT_SCALE),
|
||||
help="reasoning effort for the in-task review under --all-model (default low; cuts the review-latency tax)",
|
||||
)
|
||||
parser.add_argument(
|
||||
|
|
|
|||
File diff suppressed because one or more lines are too long
|
|
@ -185,6 +185,7 @@ Used by `commit_reviewed` for all changes to the Ouroboros repository.
|
|||
| 27 | canonical_memory_fork | If the diff touches Project/fork/execution roots, summaries, memory, or GC, does it preserve one canonical identity and distinguish authority/biography from execution-local state? Are referenced canonical artifacts promoted or retained before a child/root is collected, with missing legacy bytes represented as gaps? | critical when applicable |
|
||||
| 28 | review_artifact_continuity | If the diff changes plan, triad, scope, advisory, or acceptance evidence, are exact artifact bodies, source selectors, candidate SHA, reviewer model/profile/thread/route continuity, and all omissions retained? Continuity, transport, and coverage discrepancies are retained as typed facts beside the exact artifact bodies; a bounded or partial reviewer view must remain DEGRADED/NOT_RUN rather than PASS, and must never be implemented by discarding, blanking, or relabeling the bodies or their original cause. | critical when applicable |
|
||||
| 29 | display_identity_replay | If the diff changes routing, steering, task cards, or history replay, does it preserve the event-time human `Project › Task` presentation snapshot in both live and replay paths while keeping opaque IDs as internal/debug facts? | advisory when applicable |
|
||||
| 30 | web_design_system | If the diff touches `web/` (modules, stylesheets, `index.html`, onboarding assets): does it conform to `docs/DESIGN.md` (type scale, foreground roles, status pairs, spacing tokens) and to the engineering rules in DEVELOPMENT.md "Design System" (no new inline visual styles, values live in `web/style.css` tokens, shared components over page-local copies)? When the diff ADDS a UI control, chip, card, dialog, or visual pattern, does it reuse an existing shared frontend primitive (name it — the registry is ARCHITECTURE.md §3 "Navigation and shared UI contracts") or state in one line why none covers the need? The authoritative definitions live in those documents — check against them, do not re-derive them here. (PASS with "Not applicable" if the diff touches nothing under `web/`.) | advisory |
|
||||
|
||||
**Timeout-policy pointer for item 18 (2026-08-23):** cognitive/review waits must
|
||||
follow the layered policy in `DEVELOPMENT.md` and the timeout data-flow in
|
||||
|
|
@ -201,11 +202,12 @@ does not create another timeout constant or a second scheduler.
|
|||
- Items 6-10, 14-15, 18-20, 23, and 25-28 are conditionally critical: FAIL only when the condition applies.
|
||||
If the condition does not apply, write verdict PASS with a short reason
|
||||
(e.g. "Not applicable — no code logic change").
|
||||
- Items 11-12, 16-17, 22, 24, and 29 are advisory: FAIL produces a warning but does not
|
||||
- Items 11-12, 16-17, 22, 24, 29, and 30 are advisory: FAIL produces a warning but does not
|
||||
block. Item 22 (`cache_friendliness`) passes with "Not applicable" when the
|
||||
diff touches no prompt/context assembly. Item 24 (`perf_lifecycle`) passes with
|
||||
"Not applicable" when the diff adds/changes no endpoint, poller, subscription,
|
||||
or timer and reads no growing store.
|
||||
or timer and reads no growing store. Item 30 (`web_design_system`) passes with
|
||||
"Not applicable" when the diff touches nothing under `web/`.
|
||||
- Item 13 (self_consistency) is conditionally critical: FAIL only when the
|
||||
mismatch falls in the `Critical surface whitelist` below AND a concrete
|
||||
stale artifact is named (specific file, line, or symbol). If no whitelisted
|
||||
|
|
@ -537,7 +539,7 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete
|
|||
| 7 | extension_namespace_discipline | `type: extension` only: does the extension register its tool/route/ws-handler/ui-tab under the namespace derived from its `name` (e.g. provider-safe tool/ws names like `ext_<len>_<token>_<surface>`, route `/api/extensions/<name>/…`)? Tool and WS short names must be alphanumeric/underscore and at most 24 characters. Namespace collisions with built-in surfaces are a concrete FAIL. If the extension uses `api.send_ws_message`, are emitted event names short/provider-safe and paired with reviewed host-owned widget `subscription` components rather than arbitrary same-origin JavaScript? If the extension declares streaming UI, is it a reviewed extension route consumed by a host-owned `stream` component? If the extension owns background resources (threads, sockets, EventSource clients, subprocesses), does it register cleanup with `api.on_unload(callback)`? If the extension declares a widget render block, is it one of the host-owned schemas (`iframe`, `module`, or declarative v1: forms/actions, markdown/code, JSON/kv/table, tabs/chart, stream/subscription, progress/poll, file/gallery/media, map/calendar/kanban, group/metric/callout), with media sourced from extension routes or safe data URLs and no arbitrary same-origin JavaScript? Nested interactive group/tab children must use stable identity and one host-owned lifecycle, while `subscription.render` stays transitively passive. For non-extension skills, verdict PASS with reason "Not applicable — type != extension." | severity-driven for applicable extensions |
|
||||
| 8 | widget_module_safety | **v5.7.0+. ``kind: "module"`` widgets only.** Does the extension-supplied ``widget.js`` avoid touching ``document.cookie``, ``localStorage``, ``sessionStorage``, ``window.parent`` data, or ``fetch``/``XMLHttpRequest`` URLs OUTSIDE ``/api/extensions/<skill>/``? The host fetches reviewed ``widget.js`` through ``GET /api/extensions/<skill>/module/<entry>``, embeds the source into a sandboxed ``<iframe srcdoc sandbox="allow-scripts">`` with no ``allow-same-origin``, and injects a parent-mediated ``fetch`` bridge that rejects paths outside the owning skill route prefix. Reviewers must still confirm at the source level that the script is NOT trying to escape the sandbox via arbitrary ``postMessage`` protocols, opaque-origin storage probes, or unauthorised cross-origin fetches. Acceptable interactions: ``fetch('/api/extensions/<skill>/...')`` (through the host bridge), ``window.OuroborosWidget.fetch('/api/extensions/<skill>/...')``, and host-supplied data attributes. Mark non-module widgets and non-extension skills PASS with reason "Not applicable". | severity-driven when kind=module |
|
||||
| 9 | inject_chat_minimization | Does any use of the `inject_chat` permission have a narrow, user-facing transport purpose? The Host Service enforces token auth, skill-source attribution, rate limits, in-flight limits, fresh executable review, enablement, and explicit content-hash-bound grants. Reviewed chat transports may carry the same raw owner text as direct chat, including slash commands such as `/panic`, `/restart`, `/review`, `/evolve`, `/bg`, and `/status`; reviewers must evaluate whether the transport itself is authorized, attributable, bounded, and user-facing rather than treating slash-shaped text as automatically forbidden. A skill that accepts external inbound traffic must still show local defense-in-depth appropriate to its transport: owner/chat binding or an equivalent access rule, bounded polling/backpressure, and no unaudited broadcast to unrelated parties. Missing local defense-in-depth is a concrete FAIL for network transports. Mark PASS with reason "Not applicable" when `inject_chat` is not declared. | critical |
|
||||
| 10 | event_subscription_minimization | Are `subscribe_event` and `subscribe_events` limited to the minimum host event topics required by the skill? `chat.outbound`, `chat.typing`, `chat.photo`, `chat.video`, and `chat.document` expose owner/agent conversation data (including delivered file bytes) and require explicit justification. Wildcards, undeclared topics, or forwarding subscribed chat content to unrelated external services are concrete FAILs. Mark PASS with reason "Not applicable" when `subscribe_event` is not declared. | critical |
|
||||
| 10 | event_subscription_minimization | Are `subscribe_event` and `subscribe_events` limited to the minimum host event topics required by the skill? `chat.outbound`, `chat.typing`, `chat.photo`, `chat.video`, `chat.document`, and `chat.links` expose owner/agent conversation data (including delivered file bytes and outbound link actions) and require explicit justification. Wildcards, undeclared topics, or forwarding subscribed chat content to unrelated external services are concrete FAILs. Mark PASS with reason "Not applicable" when `subscribe_event` is not declared. | critical |
|
||||
| 11 | companion_process_safety | For `companion_process` / `supervised_task` skills: is every command declared as an argument list (not shell string), using an allowlisted runtime, with no writes outside `skill_dir` / `state_dir`, no unbounded restart loop, and cleanup on unload/panic? Does the process avoid inheriting secrets except through reviewed `env_from_settings` grants? Mark PASS with reason "Not applicable" when no long-lived process/task is declared — a transient `subprocess.run`/`subprocess.Popen` invocation of a build tool like `ffmpeg`, `ImageMagick`, or `git` inside a normal request handler is NOT a long-lived companion process and does not trigger this item (its safety belongs under items 4 / 6 / 13). | severity-driven when applicable |
|
||||
| 12 | host_token_handling | If the skill calls the Host Service API, does it use the provided `SkillToken.use_in_request()` only at request construction sites, avoid logging/serializing tokens, and keep all host-service calls on the loopback endpoint? Printing, persisting, exfiltrating, or embedding the token into user-visible output is a concrete FAIL. Mark PASS with reason "Not applicable" when the skill does not access the Host Service API. | critical |
|
||||
| 13 | error_handling | Does the skill surface actionable errors instead of swallowing exceptions, returning success on partial failure, or leaving users to inspect raw logs manually? Are retry/backoff paths bounded and purpose-specific? | advisory |
|
||||
|
|
|
|||
|
|
@ -34,6 +34,25 @@ runs. Not assumed hostile — assumed uncontrolled.
|
|||
A delegated READ-ONLY child (`mode: ask`, `access: readonly`) is not this actor. It gets no
|
||||
shell that can mutate, and it stays inside Claudexor's ordinary envelope.
|
||||
|
||||
## 2a. Stable project identity and the persistent registration (#362)
|
||||
|
||||
Fresh mutating delegated starts on an engine satisfying the workspace-root release
|
||||
contract (`CLAUDEXOR_DELEGATED_WORKSPACE_ROOT_MIN_VERSION = "3.8.1"`, the next
|
||||
compatible release carrying Claudexor PR216 after the pinned 3.8.0) register and retain
|
||||
the user's actual target project in `scope.root`, while the child's writable filesystem
|
||||
rides separately as the private snapshot in `execution.workspaceRoot`. That registration
|
||||
is the USER'S identity, not a disposable snapshot: it is marked `project_persistent`
|
||||
(`delegate_registration_policy.persistent_registration` — stable execution workspace
|
||||
plus `workspace_write`), and every retire path honours the marker — settlement, the
|
||||
orphan sweep, recovered-invocation refusals, and the pre-run refusal path. The ownership
|
||||
duty is discharged durably (`PROJECT_RETIRED` with `project_kept: true`) without
|
||||
deleting the project, and any persistent sharer makes the shared project undeletable for
|
||||
its non-persistent siblings. Durable records written before the marker existed fall back
|
||||
to the immutable stored request (`execution.workspaceRoot` + `access`) at recovery time
|
||||
(`record_persistent`); rows that carry the key are authoritative and are never
|
||||
recomputed from a live engine. Older engines keep the legacy snapshot-in-`scope.root`
|
||||
shape and retire their one-shot registration as before.
|
||||
|
||||
## 3. What Ouroboros actually controls
|
||||
|
||||
Only ADMISSION and REPORTING. Ouroboros is an HTTP client of a daemon it does not build, ship,
|
||||
|
|
|
|||
|
|
@ -203,6 +203,40 @@ not child-task cards and never prove execution by themselves.
|
|||
Codex/OpenAI comes from SVGL. Product names and marks remain the property of
|
||||
their owners.
|
||||
|
||||
### Quiz card
|
||||
|
||||
The owner quiz card (`web/modules/chat_decision.js`, `.chat-quiz-*` in
|
||||
`web/style.css`) is a chat-delivered decision surface. It is fire-and-continue
|
||||
UI: the asking task keeps working under a stated assumption, so the card must
|
||||
read correctly both as an invitation ("you can redirect me") and, after it
|
||||
settles, as a record of the path taken. Anatomy, top to bottom:
|
||||
|
||||
1. **Head** — neutral `Question` chip (`--type-meta`, neutral pair) and a
|
||||
status as dot + text. The lifecycle word family is closed:
|
||||
`Awaiting answer` (neutral dot), `Answered` (ok dot),
|
||||
`Task finished — question expired` (disabled dot), `Superseded by a retry`
|
||||
(disabled dot); an unknown state keeps a neutral dot but reads as settled
|
||||
`Closed`, never as an open invitation. No timers, no countdowns: a quiz expires only with its
|
||||
asking task.
|
||||
2. **Question** — the one primary thing: `--type-body` semibold,
|
||||
`--text-primary`.
|
||||
3. **Stake** — optional one-liner (`At stake: …`), `--type-meta`, `--text-meta`.
|
||||
4. **Options** — real owner actions: buttons with `--text-primary` labels,
|
||||
legible at rest; an optional per-option detail steps down to meta ink.
|
||||
After settlement buttons drop to `--text-disabled`; the chosen option keeps
|
||||
the ok pair. Options are capped by the shared Python↔JS constant
|
||||
(`MAX_QUIZ_OPTIONS`).
|
||||
5. **Assumption** — the signature line (`Continuing meanwhile: …`),
|
||||
`--type-meta`, `--text-meta`, separated by a hairline. While the card is
|
||||
open it names the default path; once the card settles it is the durable
|
||||
record of what the agent did without an answer. It is never dropped on
|
||||
state change.
|
||||
|
||||
The card is born on tokens even though the surrounding chat surface is not yet
|
||||
migrated: type sizes and every colour come from tokens (no new literals);
|
||||
component geometry (card min/max width, chip radius) keeps local literals like
|
||||
the rest of the chat surface until its migration pass.
|
||||
|
||||
## 6. Account group / row anatomy
|
||||
|
||||
For a repeated identity row (a connected agent account, a reviewer slot,
|
||||
|
|
@ -257,7 +291,9 @@ The scale is applied surface by surface. Migrated today:
|
|||
- `web/settings.css` (settings shell, model/effort cards, MCP cards)
|
||||
- `web/onboarding.css` (the whole first-run wizard)
|
||||
- `web/style.css` between the `design-system:migrated-begin` and
|
||||
`design-system:migrated-end` markers — harness accounts and reviewer slots
|
||||
`design-system:migrated-end` markers — harness accounts, reviewer slots,
|
||||
and the Dashboard → Updates tab (status card, one action row, collapsed
|
||||
Recovery with a single restore list)
|
||||
- the global `.muted`, `.form-section h3` and shared `.ui-status` tone rules
|
||||
|
||||
Not yet migrated: chat, skills, marketplace, widgets, logs, evolution. They are
|
||||
|
|
|
|||
|
|
@ -1104,8 +1104,10 @@ may supply that independent context; same-conversation self-review does not.
|
|||
Unavailable review is recorded as `NOT_RUN`, never silently presented as clean.
|
||||
`CONTRIBUTING.md` owns the public procedure and evidence fields.
|
||||
|
||||
`scripts/run_external_review.py --contributor` is an optional structured
|
||||
producer for the same evidence. It preserves and freezes the machine's
|
||||
`scripts/run_external_review.py --contributor` is maintainer-grade
|
||||
large-window tooling that produces structured review evidence (its scope
|
||||
reviewer's required-artifact pack is independent of diff size and can exceed
|
||||
a default install's scope window — see CONTRIBUTING for the budget shape). It preserves and freezes the machine's
|
||||
configured `api_chat` and `agent_session` triad/scope rows, then binds each row
|
||||
to its dispatched prompt receipt and observed response receipt. The shareable
|
||||
packet records exact base/head/tree/diff hashes, route/model/profile facts,
|
||||
|
|
@ -1772,7 +1774,14 @@ Before every commit, verify the following:
|
|||
- A delegating parent must not produce a clean no-tool final answer while direct
|
||||
children are still running and undecided. One bounded absorption reminder is
|
||||
allowed; after that, finalization is best-effort (`children_unabsorbed`) rather
|
||||
than clean. This is an outcome-honesty rule, not a new wait loop.
|
||||
than clean. This is an outcome-honesty rule, not a new wait loop. While the
|
||||
gate is open the delivery candidate is HELD
|
||||
(`child_absorption_or_revision_required`), never armed: the JSON-only
|
||||
delivery-control instruction must not ride the same round as the absorption
|
||||
reminder — nor a post-tool evidence change while that hold is active —
|
||||
because it would contradict the required disposition tool call. Before the
|
||||
first reminder places the hold, no disposition instruction exists yet, so
|
||||
the ordinary evidence-change arm still applies there.
|
||||
|
||||
#### Page Header Layout
|
||||
- Top-level page chrome (`renderPageHeader`, tab strips, primary actions) must sit outside the scrolling content region.
|
||||
|
|
@ -2190,12 +2199,12 @@ Before every commit, verify the following:
|
|||
#### Loop / State-Machine Changes
|
||||
- [ ] Changes to `loop.py` or other task state-machine logic include adversarial tests for malformed output, false-completion prevention, replay/log durability, and failure modes — not just the happy path.
|
||||
- [ ] Audit/checkpoint rounds must not silently reuse the normal final-answer path unless that invariant is explicitly tested and documented.
|
||||
- [ ] Keep a complete loop-local `DeliveryCandidate` once a substantive answer exists. A service round may return `keep`, or `replace` plus the complete replacement answer; allow one repair for malformed control, then preserve the prior complete answer and mark finalization degraded. A FORCED finalization (budget/round/deadline/provider/children rails) resolves an armed control purely and without retry instead: valid keep/replace is honored, anything malformed preserves the retained candidate with a typed degraded reason, and the protocol JSON itself never reaches chat or the durable result. A service notice alone does not change evidence. Owner messages, tool effects, child results, and verification receipts advance the evidence revision and require fresh delivery/acceptance binding. Finalize task-scoped service outputs/errors before host acceptance and require a complete replacement when that evidence changes; keep the `finally` path as idempotent cleanup only. This control must not bypass verification, acceptance, safety, skill-finalization, deadline, child-handoff, the unconditional `FINAL ANSWER:` latch, or the task-level answer protocol.
|
||||
- [ ] Keep a complete loop-local `DeliveryCandidate` once a substantive answer exists. A service round may return `keep`, or `replace` plus the complete replacement answer; allow one repair for malformed control, then preserve the prior complete answer and mark finalization degraded. The "malformed" class includes a protocol object embedded as a balanced trailing JSON object at the end of prose (and any fenced variant after the shared fence-strip) — never honored, never published raw while the latch is ARMED; three disclosed, test-pinned residuals: a mid-prose quotation of the literal stays prose; with the latch OFF prose with a trailing protocol object passes through the ordinary resolver as ordinary text; and a TRUNCATED (unbalanced) trailing protocol fragment passes through the forced resolver as prose even under an armed latch — a fragment is not a parseable object, and containing it would need the substring scanning the rule rejects. A FORCED finalization (budget/round/deadline/provider/children rails) resolves an armed control purely and without retry instead: valid keep/replace is honored, anything malformed preserves the retained candidate with a typed degraded reason, and the protocol JSON itself never reaches chat or the durable result. A service notice alone does not change evidence. Owner messages, tool effects, child results, and verification receipts advance the evidence revision and require fresh delivery/acceptance binding. Finalize task-scoped service outputs/errors before host acceptance and require a complete replacement when that evidence changes; keep the `finally` path as idempotent cleanup only. This control must not bypass verification, acceptance, safety, skill-finalization, deadline, child-handoff, the unconditional `FINAL ANSWER:` latch, or the task-level answer protocol.
|
||||
- [ ] Every direct child result needs an exact-hash disposition through the existing `tree_note(kind="decision")` tagged payload (`type=child_result_disposition`, child id, `integrated | irrelevant | deferred`, complete-result SHA-256; note text is rationale). One call may instead carry a `children` array of such entries (batch form): each entry is validated exactly like the single form, invalid entries are rejected individually by index while valid ones record. The typed task-tree row is the sole authority; task-result disposition fields are derived reads, never a mirrored write. The join-ledger helper alone validates lineage and current content. Stale or malformed payloads change nothing. `deferred` suppresses only the unchanged reminder and forces an honest degraded/best-effort terminal answer until the item is resolved. Natural completion WINS a late cancellation (owner decision 4=A, 2026-08-11): a child that settled its own completed result keeps it — payload, artifacts, and cost — and the cancel settles as already-settled; discarding a kept result is the parent's separate explicit `discard_child_result`. A cancelled (not completed) child still has its salvageable output preserved on the canonical drive before its bounded scratch is removed. Only a SETTLED `cancelled` status counts as a handled cancellation disposition: a child wedged in the legacy `cancel_requested` STATUS latch is intent, not outcome — it stays visible in the parent's handoff reminder as cancel-pending until custody settles it.
|
||||
- [ ] Host task acceptance is root-only. Queued/headless/scheduled roots are reviewed in `auto` and `required`; direct eligibility is the union of `outcomes.turn_has_reviewable_effects` and a typed deliverable/criterion. Ordinary read-only tool activity, pure conversation, and meta/routing controls are not reviewed, and child reviews remain advisory. Eligibility must use structured facts, never keywords (Bible P3/P5). For an eligible root under `auto|required`, agent-callable `task_acceptance_review` validates/stores evidence and optional agent disposition but makes zero reviewer calls; it returns `deferred_to_host_acceptance`, `authoritative=false`, and the evidence revision. The call itself never widens eligibility; child and `off` behavior remain unchanged.
|
||||
- [ ] Before root acceptance, atomically fence new descendants under the queue lock and prove recursive subtree quiescence from the existing task-status SSOT. Split-drive ACK, subtree, and acceptance-timing reads/writes use canonical `budget_drive_root`. Preserve the prior verdict until the replacement is recorded. A revision must explicitly reopen the fence; terminal/degraded outcomes seal it.
|
||||
- [ ] The host runs the authoritative acceptance panel once per unchanged candidate-hash/evidence-revision/fence binding. Task-acceptance actors receive one substantive call and at most two physical attempts total. Record transport status, parse status, and valid-response semantic verdict separately, with actor model/provider, role, coverage, panel id, quorum contribution, reason, enforcement impact, and binding hashes. Public task/event/UI records receive only the compact projection; full model payloads remain in private audit storage. `adaptive_quorum` applies; any contributing FAIL fails, DEGRADED abstains (the reviewer verdict vocabulary `PASS|FAIL|DEGRADED` is NOT narrowable — `_contract_valid_actors`, the deliberate-DEGRADED capsule rail and the host's core-overflow DEGRADED all depend on it), and no quorum is a terminal HOST decision. The host acceptance decision itself is written ONLY by `loop._set_acceptance_decision` and has exactly three owner-facing states — `accepted | revision_requested | finalized_unaccepted` — each with a typed `reason` from an existing structured fact; an unknown status fails closed to `finalized_unaccepted` keeping its raw token as the reason. When you add a writer, add its reason to the closed set AND check every value-keyed reader: `outcomes.derive_loop_outcome` keys the eligible-but-skipped degradation on the status+reason PAIR (`review_skipped_deadline_reserve` plus the closed forced-rail `ACCEPTANCE_BYPASS_REASONS`), and breaking that pairing is a silent false green. Forced exits stamp their typed bypass record in the common terminal recorder (`_record_forced_acceptance_bypass`) as a pure ledger write — never a fence, panel, extra round, or prompt text on a forced path, and never overwriting an existing host decision — with ONE exception (owner decision Q2A, 2026-08-10): the forced `children_unabsorbed` rail still runs the acceptance panel for an acceptance-eligible root when the subtree is quiescent, with the undispositioned-children debt included in the evidence packet; because that rail cannot take another round, a requested revision terminalizes as `finalized_unaccepted` with the typed `revision_unavailable_on_forced_rail` reason, while the process outcome stays best-effort `children_unabsorbed`. The agent may write only `agent_disposition`/`agent_rationale`, merged into the host decision, never replacing it. Clean requires PASS + solved + supported criterion evidence. Chat and Logs must use the same severity reducer, and degraded review or best-effort/degraded objective must never render as green solved. Do not add task scope review or reuse the commit gate.
|
||||
- [ ] The acceptance improvement loop is a reviewer-authored DIALOGUE (v6.74.0): obligation identity comes from the reviewer's typed `disposition_kind`/`obligation_id` (an unknown re-raise id fails closed to `new`, disclosed — never a silent fresh hash id); a re-raise reopens the row WITHOUT wiping the agent's argument (`previous_disposition`/`previous_reason`/`reopened_count` survive into the evidence catalog and the obligations clause); termination beyond a clean PASS/accepted rebuttal happens ONLY via the reviewers' quorum `dialogue_status` judgement reduced over ALL contract-valid actors (`aggregate_dialogue_status` — never `_contributing_actors`, which drops a DEGRADED slot's vote) or a real rail — no host counters, no answer/verdict hashes, no keyword gates (P5). Changes here must cover: malformed reviewer output, unknown/stale `obligation_id` on a re_raise, partial panel failure, multi-slot dialogue-status disagreement (the reducer's precedence), replay/restart durability of obligation rows, false completion, and the backward-compatible default when the new fields are absent.
|
||||
- [ ] The host buys one authoritative acceptance panel per PAID IDENTITY: `sha256(candidate_hash + the sorted set of nonempty (obligation_id, disposition, sha256(reason)) tuples)` (owner ratification 2026-08-30, replacing the earlier candidate-hash/evidence-revision/fence binding rule). Only two things mint a new paid panel — a changed candidate answer, or a new nonempty obligation disposition; an empty disposition reason hashes to `""` and buys nothing (mirroring `commit_gate.compute_rebuttal_sha256`). The evidence revision must NOT mint a paid cycle — every cosmetic tool call moves it — and remains stale-packet detection for the supersede paths. A resubmit with an unchanged paid identity replays the recorded verdict for FREE, must NOT re-enter the improvement capsule, and terminalizes with the typed `identical_acceptance_refused` reason. Task-acceptance actors receive one substantive call and at most two physical attempts total. Record transport status, parse status, and valid-response semantic verdict separately, with actor model/provider, role, coverage, panel id, quorum contribution, reason, enforcement impact, and binding hashes. Public task/event/UI records receive only the compact projection; full model payloads remain in private audit storage. `adaptive_quorum` applies; any contributing FAIL fails, DEGRADED abstains (the reviewer verdict vocabulary `PASS|FAIL|DEGRADED` is NOT narrowable — `_contract_valid_actors`, the deliberate-DEGRADED capsule rail and the host's core-overflow DEGRADED all depend on it), and no quorum is a terminal HOST decision. The host acceptance decision itself is written ONLY by `acceptance_dialogue._set_acceptance_decision` (re-exported from `loop`) and has exactly three owner-facing states — `accepted | revision_requested | finalized_unaccepted` — each with a typed `reason` from an existing structured fact; an unknown status fails closed to `finalized_unaccepted` keeping its raw token as the reason. When you add a writer, add its reason to the closed set AND check every value-keyed reader: `outcomes.derive_loop_outcome` keys the eligible-but-skipped degradation on the status+reason PAIR (`review_skipped_deadline_reserve` plus the closed forced-rail `ACCEPTANCE_BYPASS_REASONS`), and the BLOCKED objective terminal on `_ACCEPTANCE_BLOCKED_TERMINAL_REASONS` (`review_cycles_exhausted`, `identical_acceptance_refused`); breaking either pairing is a silent false green. Forced exits stamp their typed bypass record in the common terminal recorder (`_record_forced_acceptance_bypass`) as a pure ledger write — never a fence, panel, extra round, or prompt text on a forced path, and never overwriting an existing host decision — with ONE exception (owner decision Q2A, 2026-08-10): the forced `children_unabsorbed` rail still runs the acceptance panel for an acceptance-eligible root when the subtree is quiescent, with the undispositioned-children debt included in the evidence packet; because that rail cannot take another round, a requested revision terminalizes as `finalized_unaccepted` with the typed `revision_unavailable_on_forced_rail` reason, while the process outcome stays best-effort `children_unabsorbed`. The agent may write only `agent_disposition`/`agent_rationale`, merged into the host decision, never replacing it. Clean requires PASS + solved + supported criterion evidence. Chat and Logs must use the same severity reducer, and degraded review or best-effort/degraded objective must never render as green solved. Do not add task scope review or reuse the commit gate.
|
||||
- [ ] The acceptance improvement loop is a reviewer-authored DIALOGUE (v6.74.0): obligation identity comes from the reviewer's typed `disposition_kind`/`obligation_id` (an unknown re-raise id fails closed to `new`, disclosed — never a silent fresh hash id); a re-raise reopens the row WITHOUT wiping the agent's argument (`previous_disposition`/`previous_reason`/`reopened_count` survive into the evidence catalog and the obligations clause); termination beyond a clean PASS/accepted rebuttal happens ONLY via the reviewers' `dialogue_status` judgement (`aggregate_dialogue_status`) or a real rail — no host counters, no keyword gates (P5). That reducer is now read over `_contributing_actors` (owner ratification 2026-08-30, replacing the earlier widening to all contract-valid actors: a slot whose verdict did not reach the aggregate does not steer the loop either), and majority voting stays REJECTED — ONE contributing reviewer may hold the loop open, but only WITH MATERIAL: a `continue_actionable` vote counts only when the same response carries a concrete finding or a completion_coach, otherwise it is disclosed as `continue_without_findings` and abstains. Missing/invalid votes abstain too (`abstain_invalid`) and NEVER default to continue; a single well-formed terminal vote ends the dialogue; ZERO well-formed votes reduce to the typed `inconclusive`, which is not reviewer vocabulary (`DIALOGUE_STATUS_VALUES` is unchanged), grants the dialogue no authority, and falls through to the existing DEGRADED / no-capsule / exhaustion terminals — never a host-minted `stable_disagreement` and never another paid round. The reviewer verdict vocabulary `PASS|FAIL|DEGRADED` remains NOT narrowable. Changes here must cover: malformed reviewer output, unknown/stale `obligation_id` on a re_raise, partial panel failure, multi-slot dialogue-status disagreement (the reducer's precedence), replay/restart durability of obligation rows, false completion, and the backward-compatible default when the new fields are absent.
|
||||
- [ ] An explicit `max_improvement_passes` binds under every legacy policy. Without one, the shared review-cycle cap binds under EVERY policy, Required+Blocking included: `OUROBOROS_REVIEW_MAX_CYCLES` (SSOT `ouroboros/review_cycles.py`; string, default `"2"`, `"unlimited"` = no local count cap) gives `improvement passes = cycles − 1`; the retired `OUROBOROS_ACCEPTANCE_MAX_IMPROVEMENT_PASSES` is migrated into the shared key at settings load and never binds at runtime.
|
||||
|
||||
#### Cognitive Artifact Integrity
|
||||
|
|
@ -2284,6 +2293,24 @@ Before every commit, verify the following:
|
|||
Guard and handler must receive identical argv. Do not rewrite explicit paths,
|
||||
versioned interpreters, shell bodies, or remote execution, and never install a
|
||||
dependency in response to `ModuleNotFoundError`.
|
||||
- Resolve bare Node (`node`/`nodejs`; Windows launcher suffixes normalize for
|
||||
matching) for the same four surfaces, but once AFTER the dispatch gates: the
|
||||
node health check is an execution probe of an argv[0]-steered candidate, and
|
||||
pre-guard it would run a planted `PATH` shim before the fences refuse the
|
||||
call. The ladder is PATH-first with the probe — a healthy `PATH` node is a
|
||||
byte-identical no-op in argv and child env — and the bundled runtime
|
||||
substitutes only when the `PATH` candidate is missing or probe-dead: rewrite
|
||||
only `node`/`nodejs` argv[0]; npm/npx/pnpm/yarn/corepack and `sh`/`bash`
|
||||
bodies naming a family token get only the attested child-env `PATH` prepend
|
||||
(a formula with a rewritten absolute shebang is a disclosed residual). Guards
|
||||
inspect the original argv; the substitution stays in the same interpreter
|
||||
family and reaches the handler through the per-call attestation, which
|
||||
`verify_and_record` uses to execute the resolved argv while `check` keeps the
|
||||
original receipt-identity text. A non-local executor backend is never touched
|
||||
(no host path leaks into a container; a local executor continues the ladder),
|
||||
explicit paths and versioned names are never rewritten, and node never
|
||||
pre-blocks: with no usable runtime the argv runs as written and fails
|
||||
honestly with the probe facts disclosed.
|
||||
- Skill Review ordinals and provenance stay in `review_job.json` and the
|
||||
append-only `review_history.jsonl`: allocate under the lifecycle lock, consume
|
||||
a round only after actual start, write one terminal row per `job_id`, and
|
||||
|
|
|
|||
|
|
@ -46,24 +46,24 @@
|
|||
<article>
|
||||
<p class="download-platform">macOS 12+ · Apple silicon only</p>
|
||||
<h2>macOS</h2>
|
||||
<a data-release-download="macos-arm64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4.dmg">Download for macOS (.dmg)</a>
|
||||
<a data-release-download="macos-arm64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5.dmg">Download for macOS (.dmg)</a>
|
||||
<p>Open the DMG, drag <code>Ouroboros.app</code> to Applications, then launch it from Applications.</p>
|
||||
</article>
|
||||
<article>
|
||||
<p class="download-platform">Windows x64</p>
|
||||
<h2>Windows</h2>
|
||||
<a data-release-download="windows-x64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-windows-x64.zip">Download for Windows (.zip)</a>
|
||||
<a data-release-download="windows-x64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-windows-x64.zip">Download for Windows (.zip)</a>
|
||||
<p>Extract the ZIP, open the <code>Ouroboros</code> folder, and run <code>Ouroboros.exe</code>.</p>
|
||||
</article>
|
||||
<article id="linux">
|
||||
<p class="download-platform">Linux x86_64</p>
|
||||
<h2>Linux</h2>
|
||||
<div class="download-actions">
|
||||
<a data-release-download="linux-deb-amd64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros_6.113.4_amd64.deb">Debian / Ubuntu / Astra (.deb)</a>
|
||||
<a data-release-download="linux-rpm-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros-6.113.4-1.x86_64.rpm">Fedora / RHEL (.rpm)</a>
|
||||
<a data-release-download="linux-rpm-red80-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/ouroboros-6.113.4-1.red80.x86_64.rpm">RED OS 8 (.rpm)</a>
|
||||
<a data-release-download="linux-appimage-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-linux-x86_64.AppImage">Portable AppImage</a>
|
||||
<a data-release-download="linux-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4-linux-x86_64.tar.gz">tar.gz archive</a>
|
||||
<a data-release-download="linux-deb-amd64" class="btn btn-primary" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros_6.113.5_amd64.deb">Debian / Ubuntu / Astra (.deb)</a>
|
||||
<a data-release-download="linux-rpm-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros-6.113.5-1.x86_64.rpm">Fedora / RHEL (.rpm)</a>
|
||||
<a data-release-download="linux-rpm-red80-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/ouroboros-6.113.5-1.red80.x86_64.rpm">RED OS 8 (.rpm)</a>
|
||||
<a data-release-download="linux-appimage-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-linux-x86_64.AppImage">Portable AppImage</a>
|
||||
<a data-release-download="linux-x86_64" class="btn btn-quiet" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5-linux-x86_64.tar.gz">tar.gz archive</a>
|
||||
</div>
|
||||
<p>Prefer the package for your distribution. Use the AppImage or archive on other distributions; Git must already be installed.</p>
|
||||
</article>
|
||||
|
|
@ -74,7 +74,7 @@
|
|||
<section class="install-primary" aria-labelledby="macos-quick-start">
|
||||
<h2 id="macos-quick-start">macOS quick start</h2>
|
||||
<ol class="install-steps">
|
||||
<li>Click <a data-release-download="macos-arm64" href="https://github.com/razzant/ouroboros/releases/download/v6.113.4/Ouroboros-6.113.4.dmg">Download for macOS (.dmg)</a>.</li>
|
||||
<li>Click <a data-release-download="macos-arm64" href="https://github.com/razzant/ouroboros/releases/download/v6.113.5/Ouroboros-6.113.5.dmg">Download for macOS (.dmg)</a>.</li>
|
||||
<li>Open the DMG and drag <code>Ouroboros.app</code> onto the <strong>Applications</strong> shortcut.</li>
|
||||
<li>Open Ouroboros from Applications. If Gatekeeper asks, right-click the app and choose <strong>Open</strong>.</li>
|
||||
</ol>
|
||||
|
|
|
|||
82
launcher.py
82
launcher.py
|
|
@ -2,6 +2,7 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
|
|
@ -12,6 +13,8 @@ import sys
|
|||
import tempfile
|
||||
import threading
|
||||
import time
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
import webbrowser
|
||||
from logging.handlers import RotatingFileHandler
|
||||
from typing import Optional
|
||||
|
|
@ -1009,7 +1012,7 @@ def _headless_signal_handler(signum, frame) -> None:
|
|||
_shutdown_event.set()
|
||||
|
||||
|
||||
def _open_browser_detached(url: str) -> threading.Thread:
|
||||
def _open_browser_detached(url: str, outcome: Optional[list] = None) -> threading.Thread:
|
||||
"""Open the default browser without ever blocking the caller.
|
||||
|
||||
`webbrowser.open` waits for the child on a stdlib-resolved console
|
||||
|
|
@ -1019,6 +1022,10 @@ def _open_browser_detached(url: str) -> threading.Thread:
|
|||
so a short-lived caller (the already-running notice) can bound-join it
|
||||
before process exit would kill the daemon thread under the opener.
|
||||
|
||||
An ``outcome`` list, when given, receives exactly one entry — True/False
|
||||
from ``webbrowser.open`` or the raised exception — so a bounded-join
|
||||
caller (the desktop bridge) can report failure honestly.
|
||||
|
||||
DELIBERATE (owner-approved): the opened browser is the USER'S own
|
||||
application, intentionally outside process custody and launcher teardown —
|
||||
the Emergency-Stop invariant governs the AGENT'S tree, and killing the
|
||||
|
|
@ -1027,9 +1034,12 @@ def _open_browser_detached(url: str) -> threading.Thread:
|
|||
"""
|
||||
def _open() -> None:
|
||||
try:
|
||||
webbrowser.open(url)
|
||||
except Exception:
|
||||
result: object = bool(webbrowser.open(url))
|
||||
except Exception as exc:
|
||||
result = exc
|
||||
log.info("Could not open the default browser for %s", url, exc_info=True)
|
||||
if outcome is not None:
|
||||
outcome.append(result)
|
||||
|
||||
thread = threading.Thread(target=_open, name="ouroboros-open-browser", daemon=True)
|
||||
thread.start()
|
||||
|
|
@ -1325,8 +1335,6 @@ def main():
|
|||
Shared SSOT for both the download-to-Downloads and open-in-default-app
|
||||
bridge methods so the loopback guard cannot drift between them.
|
||||
"""
|
||||
import urllib.parse
|
||||
|
||||
full_url = urllib.parse.urljoin(f"http://127.0.0.1:{actual_port}", str(raw_url or ""))
|
||||
parsed = urllib.parse.urlparse(full_url)
|
||||
if parsed.scheme != "http":
|
||||
|
|
@ -1335,16 +1343,15 @@ def main():
|
|||
raise ValueError("desktop file access is limited to the local Ouroboros server")
|
||||
if parsed.port != actual_port:
|
||||
raise ValueError("file URL port must match the local Ouroboros server")
|
||||
if parsed.path != "/api/files/download" and not parsed.path.startswith("/api/extensions/"):
|
||||
raise ValueError("file URL path must be /api/files/download or /api/extensions/<skill>/...")
|
||||
if parsed.path != "/api/files/download" and not parsed.path.startswith(("/api/extensions/", "/api/tasks/")):
|
||||
raise ValueError("file URL path must be /api/files/download, /api/extensions/<skill>/... or /api/tasks/...")
|
||||
return full_url
|
||||
|
||||
def _unique_bridge_target(directory: pathlib.Path, filename: str) -> pathlib.Path:
|
||||
safe_name = pathlib.Path(str(filename or "download")).name or "download"
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
target = directory / safe_name
|
||||
stem = target.stem
|
||||
suffix = target.suffix
|
||||
stem, suffix = target.stem, target.suffix
|
||||
counter = 1
|
||||
while target.exists():
|
||||
target = directory / f"{stem}-{counter}{suffix}"
|
||||
|
|
@ -1352,46 +1359,31 @@ def main():
|
|||
return target
|
||||
|
||||
def _fetch_bridge_url_to(full_url: str, target: pathlib.Path) -> None:
|
||||
import urllib.request
|
||||
|
||||
with urllib.request.urlopen(full_url, timeout=60) as resp: # noqa: S310 - localhost validated above
|
||||
with target.open("wb") as fh:
|
||||
shutil.copyfileobj(resp, fh)
|
||||
with urllib.request.urlopen(full_url, timeout=60) as resp, target.open("wb") as fh: # noqa: S310 - localhost validated above
|
||||
shutil.copyfileobj(resp, fh)
|
||||
|
||||
class MainApi:
|
||||
@staticmethod
|
||||
def _native_confirm(title: str, message: str) -> bool:
|
||||
return bool(_webview_window and _webview_window.create_confirmation_dialog(title, message))
|
||||
|
||||
def request_runtime_mode_change(self, mode: str) -> dict:
|
||||
try:
|
||||
return _request_runtime_mode_change(
|
||||
mode,
|
||||
lambda title, message: bool(
|
||||
_webview_window and _webview_window.create_confirmation_dialog(title, message)
|
||||
),
|
||||
)
|
||||
return _request_runtime_mode_change(mode, self._native_confirm)
|
||||
except Exception as exc:
|
||||
log.warning("Runtime mode native confirmation failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": f"Native confirmation failed: {exc}"}
|
||||
|
||||
def request_auto_grant_reviewed_skills_change(self, enabled: bool) -> dict:
|
||||
try:
|
||||
return _request_auto_grant_reviewed_skills_change(
|
||||
bool(enabled),
|
||||
lambda title, message: bool(
|
||||
_webview_window and _webview_window.create_confirmation_dialog(title, message)
|
||||
),
|
||||
)
|
||||
return _request_auto_grant_reviewed_skills_change(bool(enabled), self._native_confirm)
|
||||
except Exception as exc:
|
||||
log.warning("Reviewed-skill auto-grant confirmation failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": f"Native confirmation failed: {exc}"}
|
||||
|
||||
def request_skill_key_grant(self, skill: str, keys: list) -> dict:
|
||||
try:
|
||||
return _request_skill_key_grant(
|
||||
skill,
|
||||
keys,
|
||||
lambda title, message: bool(
|
||||
_webview_window and _webview_window.create_confirmation_dialog(title, message)
|
||||
),
|
||||
)
|
||||
return _request_skill_key_grant(skill, keys, self._native_confirm)
|
||||
except Exception as exc:
|
||||
log.warning("Skill grant native confirmation failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": f"Native confirmation failed: {exc}"}
|
||||
|
|
@ -1408,6 +1400,30 @@ def main():
|
|||
log.warning("Desktop file download failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": str(exc)}
|
||||
|
||||
def open_external_url(self, url: str) -> dict:
|
||||
try:
|
||||
raw = str(url or "")
|
||||
if not raw.lower().startswith(("http://", "https://", "mailto:")):
|
||||
return {"ok": False, "error": "Only absolute http://, https:// or mailto: links can be opened."}
|
||||
outcome: list = []
|
||||
# Bounded join: settled failure reported honestly; still-running stays detached.
|
||||
_open_browser_detached(raw, outcome).join(timeout=3.0)
|
||||
if outcome and outcome[0] is not True:
|
||||
return {"ok": False, "error": f"The default browser could not be opened: {outcome[0] or 'no handler found'}"}
|
||||
return {"ok": True}
|
||||
except Exception as exc:
|
||||
log.warning("Desktop external-URL open failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": str(exc)}
|
||||
|
||||
def save_bytes_to_downloads(self, filename: str, b64: str) -> dict:
|
||||
try:
|
||||
target = _unique_bridge_target(pathlib.Path.home() / "Downloads", filename)
|
||||
target.write_bytes(base64.b64decode(str(b64 or ""), validate=True))
|
||||
return {"ok": True, "path": str(target)}
|
||||
except Exception as exc:
|
||||
log.warning("Desktop save-to-Downloads failed: %s", exc, exc_info=True)
|
||||
return {"ok": False, "error": str(exc)}
|
||||
|
||||
def open_file_with_default_app(self, url: str, filename: str) -> dict:
|
||||
"""Open a delivered file in the OS default app (external window).
|
||||
|
||||
|
|
|
|||
|
|
@ -626,6 +626,18 @@ def verification_receipt_ledger_row(receipt: Dict[str, Any]) -> Dict[str, Any]:
|
|||
"expected_match": str(receipt.get("expected_match") or "substring"),
|
||||
"matched": receipt.get("matched"),
|
||||
"returncode": receipt.get("returncode"),
|
||||
# Disclosure keys (node-runtime sprint, D6/R4) — carried ONLY when the
|
||||
# receipt has them, so historical rows and non-run kinds stay
|
||||
# byte-identical. `duration_ms`: how long the check's process lived.
|
||||
# `signal`: POSIX signal name of a killed check (returncode < 0) —
|
||||
# a 9ms SIGKILL is distinguishable from an honest red. `resolved_runtime`:
|
||||
# the physical executable that ran when the interpreter resolver
|
||||
# substituted one; ABSENT means the recorded `check` argv executed as
|
||||
# written. Disclosure only — receipt identity/reconciliation (the shared
|
||||
# `receipt_identity_projection` above) reads none of these.
|
||||
**({"duration_ms": receipt.get("duration_ms")} if receipt.get("duration_ms") is not None else {}),
|
||||
**({"signal": str(receipt.get("signal") or "")} if receipt.get("signal") else {}),
|
||||
**({"resolved_runtime": truncate_review_artifact(receipt.get("resolved_runtime"), limit=300)} if receipt.get("resolved_runtime") else {}),
|
||||
"summary": truncate_review_artifact(receipt.get("summary"), limit=300),
|
||||
# C: after-only artifact-lifecycle flag (a check that built then deleted a
|
||||
# declared deliverable). Bounded through the SHARED disclosed-list projection —
|
||||
|
|
|
|||
|
|
@ -159,11 +159,39 @@ _POLICY_DENIAL_STATUSES = frozenset({
|
|||
# DELIBERATELY demotes them to a non-degrading "cosmetic" bucket (still recorded
|
||||
# on the execution axis for monitoring) because the owner accepted that an
|
||||
# ignored shell failure belongs on the LLM-review/objective axis, not the
|
||||
# execution axis. `timeout` is intentionally EXCLUDED — a stuck/aborted command
|
||||
# is a real failure. Structural status/tool-name partition, never content
|
||||
# matching (Bible P5).
|
||||
# execution axis. TWO exclusions keep the demotion honest, symmetric in intent:
|
||||
# `timeout` (never in the set — a stuck/aborted command is a real failure) and
|
||||
# SIGNAL DEATH (D7, node-runtime sprint: a process killed by a signal —
|
||||
# `_is_signal_death` below — is excluded at the branch, because a kernel
|
||||
# CODESIGNING kill, an OOM kill, or an external `kill -9` is an
|
||||
# execution-degrading fact, not an ignorable probe exit; the 43-SIGKILL macOS
|
||||
# incident hid exactly here as "cosmetic"). Structural status/tool-name/typed-
|
||||
# meta partition, never content matching (Bible P5).
|
||||
_NON_BLOCKING_RECOVERABLE_STATUSES = frozenset({"non_zero_exit", "shell_error"})
|
||||
_COSMETIC_TOOL_NAMES = frozenset({"run_command", "run_script"})
|
||||
|
||||
|
||||
def _is_signal_death(item: Dict[str, Any]) -> bool:
|
||||
"""TYPED signal-death test for one tool-call record (D7): the child died to
|
||||
a signal when the record's meta shows a NEGATIVE exit code or a signal name.
|
||||
|
||||
Reads the typed ``exit_code``/``signal`` fields that the process-tool
|
||||
handler publishes structurally (loop_tool_execution merges them into
|
||||
result_meta); for LEGACY records the same keys were regex-harvested from
|
||||
the rendered text, so the read covers both generations. A record carrying
|
||||
NEITHER key cannot prove signal death and stays in the cosmetic bucket.
|
||||
|
||||
POSIX semantics — declared residual, not faked portability: on Windows a
|
||||
killed process reports a large POSITIVE exit code (e.g. 0xC0000005) and no
|
||||
signal, so this partition cannot see a Windows kill; such a record remains
|
||||
cosmetic."""
|
||||
if str(item.get("signal") or "").strip():
|
||||
return True
|
||||
exit_code = item.get("exit_code")
|
||||
try:
|
||||
return exit_code is not None and int(exit_code) < 0
|
||||
except (TypeError, ValueError):
|
||||
return False
|
||||
# A2: an UNRECOVERED access-policy block (resource_policy_blocked /
|
||||
# resource_constraint_blocked) on a READ-ONLY exploratory tool — e.g. a
|
||||
# read_file/search_code/query_code refused by the resource policy — is honest
|
||||
|
|
@ -238,6 +266,12 @@ def _tool_error_record(item: Dict[str, Any], *, recovered_by: int | None = None)
|
|||
"signal": item.get("signal"),
|
||||
"result": str(item.get("result") or "")[:500],
|
||||
}
|
||||
# T11: forensic disclosure for the signal-death class - a millisecond
|
||||
# lifetime names a kernel kill, and the attested runtime separates a broken
|
||||
# PATH node from the bundled substitute. Absent keys stay absent.
|
||||
for extra in ("duration_ms", "resolved_runtime"):
|
||||
if item.get(extra) is not None:
|
||||
record[extra] = item.get(extra)
|
||||
if recovered_by is not None:
|
||||
record["recovered_by_call_index"] = recovered_by
|
||||
return record
|
||||
|
|
@ -337,8 +371,15 @@ def _classify_tool_errors(llm_trace: Dict[str, Any]) -> Dict[str, List[Dict[str,
|
|||
if recovered_by is not None:
|
||||
recovered_items.append(_tool_error_record(item, recovered_by=recovered_by))
|
||||
continue
|
||||
if status in _NON_BLOCKING_RECOVERABLE_STATUSES and tool in _COSMETIC_TOOL_NAMES:
|
||||
# Unrecovered run_command/run_script non-zero exit: cosmetic, not degrading.
|
||||
if (
|
||||
status in _NON_BLOCKING_RECOVERABLE_STATUSES
|
||||
and tool in _COSMETIC_TOOL_NAMES
|
||||
and not _is_signal_death(item)
|
||||
):
|
||||
# Unrecovered run_command/run_script non-zero exit: cosmetic, not
|
||||
# degrading. Signal death (negative exit_code / signal name in the
|
||||
# typed meta) falls THROUGH to the unresolved bucket — symmetric
|
||||
# with the timeout exclusion above (D7).
|
||||
cosmetic_items.append(_tool_error_record(item))
|
||||
continue
|
||||
if status in _POLICY_DENIAL_STATUSES:
|
||||
|
|
|
|||
|
|
@ -1,11 +1,16 @@
|
|||
"""In-process validated-rows memo and render cache for ledger display reads.
|
||||
"""In-process read acceleration for the usage ledger, beside the substrate.
|
||||
|
||||
Extracted from ``usage_accounting.py`` at the perf2 P1 render-cache round for
|
||||
the same reason ``_usage_rows.py`` exists: the accounting module sits at the
|
||||
hard module-size gate and this layer is a self-contained seam. It is the READ
|
||||
acceleration for display projections only — write paths keep their own full
|
||||
locked reads — and it deliberately lives beside, not inside,
|
||||
``usage_ledger.py``: the substrate stays cache-ignorant.
|
||||
hard module-size gate and this layer is a self-contained seam. It carries the
|
||||
ledger's two in-process caches — the validated-rows memo + render cache for
|
||||
display projections, and the in-lock warm read cache the monetary write paths
|
||||
use (razzant/ouroboros#129) — and it deliberately lives beside, not inside,
|
||||
``usage_ledger.py``: the substrate stays cache-ignorant. In every case the
|
||||
full ``_read_records_locked`` replay remains the authority and the sole owner
|
||||
of quarantine; a cache can only ever change the COST of a read, never its
|
||||
result, because ``_read_new_records_locked`` re-stats the file under the held
|
||||
lock and refuses to resume on any doubt.
|
||||
|
||||
``usage_accounting`` re-binds every name here, and the implementation resolves
|
||||
the substrate (``_locked``, ``_read_records_locked``, ...) through the
|
||||
|
|
@ -15,7 +20,9 @@ exactly as when the code was inline.
|
|||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import collections
|
||||
import copy
|
||||
import logging
|
||||
import pathlib
|
||||
import threading
|
||||
from dataclasses import dataclass, field
|
||||
|
|
@ -23,6 +30,8 @@ from typing import Any, Callable, Dict, Tuple
|
|||
|
||||
from ouroboros.usage_ledger import QUARANTINE_REL, LedgerResumeState
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _ua():
|
||||
"""The accounting namespace, resolved lazily (import cycle + test pins)."""
|
||||
|
|
@ -54,10 +63,12 @@ class _LedgerRowsMemo:
|
|||
|
||||
# Read-side memo per RESOLVED drive root. Populated and advanced only under the
|
||||
# cross-process ledger lock; the module lock guards the dict itself. Write paths
|
||||
# (reserve/_transition/settle/import) never touch it — their own full locked
|
||||
# reads stay the monetary authority, and the stat + seq-continuity check on the
|
||||
# next read is what makes a stale memo impossible to serve, so correctness never
|
||||
# depends on any writer remembering to invalidate.
|
||||
# (reserve/_transition/settle/import) never touch it — they read through their
|
||||
# own in-lock cache below (full ordered records, which seq assignment and
|
||||
# whole-history append validation need; this memo keeps only final rows), and
|
||||
# the stat + seq-continuity check on the next read is what makes a stale memo
|
||||
# impossible to serve, so correctness never depends on any writer remembering
|
||||
# to invalidate.
|
||||
_ROWS_MEMO: Dict[str, _LedgerRowsMemo] = {}
|
||||
_ROWS_MEMO_LOCK = threading.Lock()
|
||||
|
||||
|
|
@ -161,3 +172,55 @@ def _render_cached(
|
|||
if _ROWS_MEMO.get(key) is memo and memo.generation == generation:
|
||||
memo.renders[full_key] = copy.deepcopy(result)
|
||||
return result
|
||||
|
||||
|
||||
# razzant/ouroboros#129: the in-lock write paths (reserve/settle/_transition/
|
||||
# release/legacy-import) each did a full parse+validate of the whole ledger
|
||||
# under the 45s monetary flock, and the file grows unboundedly. This is their
|
||||
# per-process warm cache of the last validated read per drive root: the next
|
||||
# in-lock read parses only the bytes appended since. It is distinct from
|
||||
# ``_ROWS_MEMO`` because writers need the FULL ordered records list (seq
|
||||
# assignment + whole-history append validation), not just final rows. Rows are
|
||||
# shared read-only snapshots, same as the memo's.
|
||||
_LEDGER_READ_CACHE: "collections.OrderedDict[str, Tuple[LedgerResumeState, list]]" = (
|
||||
collections.OrderedDict()
|
||||
)
|
||||
_LEDGER_READ_CACHE_LOCK = threading.Lock()
|
||||
_LEDGER_READ_CACHE_MAX_ROOTS = 8
|
||||
|
||||
|
||||
def _ledger_cache_put(key: str, value: "Tuple[LedgerResumeState, list]") -> None:
|
||||
with _LEDGER_READ_CACHE_LOCK:
|
||||
_LEDGER_READ_CACHE[key] = value
|
||||
_LEDGER_READ_CACHE.move_to_end(key)
|
||||
while len(_LEDGER_READ_CACHE) > _LEDGER_READ_CACHE_MAX_ROOTS:
|
||||
_LEDGER_READ_CACHE.popitem(last=False)
|
||||
|
||||
|
||||
def _read_records_locked_cached(root: pathlib.Path) -> list:
|
||||
"""``_read_records_locked`` with an incremental warm path. Call under the
|
||||
held ledger lock (same contract as ``_read_records_locked``)."""
|
||||
ua = _ua()
|
||||
key = str(pathlib.Path(root).resolve(strict=False)) # one slot per physical root
|
||||
with _LEDGER_READ_CACHE_LOCK:
|
||||
cached = _LEDGER_READ_CACHE.get(key)
|
||||
if cached is not None:
|
||||
resume, rows = cached
|
||||
try:
|
||||
delta = ua._read_new_records_locked(root, resume)
|
||||
except Exception: # noqa: BLE001 — any doubt = fall back to the full read
|
||||
delta = None
|
||||
if delta is not None:
|
||||
new_rows, new_resume = delta
|
||||
merged = rows if not new_rows else [*rows, *new_rows]
|
||||
_ledger_cache_put(key, (new_resume, merged))
|
||||
return list(merged)
|
||||
records = ua._read_records_locked(root)
|
||||
try:
|
||||
resume = ua._ledger_resume_state(root, records)
|
||||
_ledger_cache_put(key, (resume, list(records)))
|
||||
except Exception: # noqa: BLE001 — caching is best-effort; correctness is the full read
|
||||
log.debug("ledger read-cache seed failed for %s", key, exc_info=True)
|
||||
with _LEDGER_READ_CACHE_LOCK:
|
||||
_LEDGER_READ_CACHE.pop(key, None)
|
||||
return records
|
||||
|
|
|
|||
|
|
@ -113,6 +113,36 @@ class Env:
|
|||
return (self.drive_root / safe_relpath(rel)).resolve()
|
||||
|
||||
|
||||
def _emit_budget_pause_checkpoint(
|
||||
event_queue: Any,
|
||||
drive_logs: pathlib.Path | None,
|
||||
task_id: str,
|
||||
resource_limit: Dict[str, Any],
|
||||
) -> None:
|
||||
"""Publish the owner-visible budget-pause checkpoint on the registered path.
|
||||
|
||||
The enveloped log_event route is the ONE registered checkpoint channel
|
||||
(_handle_log_event persists task_checkpoint rows and pushes them live);
|
||||
a bare task_checkpoint on _pending_events has no handler and died as
|
||||
unknown_worker_event, silently losing the owner-visible toast.
|
||||
"""
|
||||
checkpoint = {
|
||||
"type": "task_checkpoint",
|
||||
"task_id": task_id,
|
||||
"checkpoint_kind": "budget_scope_paused",
|
||||
"owner_visible": True,
|
||||
"toast_once": f"{task_id}:budget-paused:{resource_limit['scope']}",
|
||||
**resource_limit,
|
||||
}
|
||||
if event_queue is not None:
|
||||
emit_log_event(event_queue, checkpoint)
|
||||
elif drive_logs:
|
||||
try:
|
||||
append_jsonl(drive_logs / "events.jsonl", {"ts": utc_now_iso(), **checkpoint})
|
||||
except Exception:
|
||||
log.debug("budget-pause checkpoint append failed", exc_info=True)
|
||||
|
||||
|
||||
class OuroborosAgent:
|
||||
"""Per-worker agent instance; long-term state lives on Drive."""
|
||||
|
||||
|
|
@ -688,11 +718,15 @@ class OuroborosAgent:
|
|||
|
||||
def handle_task(self, task: Dict[str, Any]) -> List[Dict[str, Any]]:
|
||||
"""Run one task under the root/subtree monetary attribution scope."""
|
||||
# Hot-reload settings so UI changes affect the next task without restart.
|
||||
try:
|
||||
subagent_runtime.apply_task_start_settings()
|
||||
except Exception:
|
||||
pass
|
||||
# A reused worker agent still carries the PREVIOUS task's chat binding;
|
||||
# events emitted before _handle_task_scoped rebinds it would be
|
||||
# addressed to the old thread. No binding lets the supervisor stamp
|
||||
# the right chat from the RUNNING row by task_id (log_addressing).
|
||||
self._current_chat_id = None
|
||||
# Hot-reload settings so UI changes affect the next task without
|
||||
# restart; a failed reload is disclosed loudly, not swallowed (#285).
|
||||
subagent_runtime.apply_task_start_settings_or_disclose(
|
||||
str(task.get("id") or ""), self._emit_live_log)
|
||||
|
||||
from ouroboros.usage_accounting import UsageScope, usage_scope
|
||||
|
||||
|
|
@ -976,14 +1010,9 @@ class OuroborosAgent:
|
|||
"tool_calls": [],
|
||||
"resource_limit": resource_limit,
|
||||
}
|
||||
self._pending_events.append({
|
||||
"type": "task_checkpoint",
|
||||
"task_id": task_id,
|
||||
"checkpoint_kind": "budget_scope_paused",
|
||||
"owner_visible": True,
|
||||
"toast_once": f"{task_id}:budget-paused:{resource_limit['scope']}",
|
||||
**resource_limit,
|
||||
})
|
||||
_emit_budget_pause_checkpoint(
|
||||
self._event_queue, drive_logs, task_id, resource_limit
|
||||
)
|
||||
emit_task_results(
|
||||
self.env,
|
||||
self.memory,
|
||||
|
|
|
|||
|
|
@ -103,7 +103,8 @@ def dispatch_executor_note(decision: Optional[SubagentExecutorResolution],
|
|||
"co-building around a $0 run. If your run asks a question (delegate_wait "
|
||||
"returns waiting_on_user), answer it from the task context with "
|
||||
"delegate_answer; a question above your authority — money, scope, external "
|
||||
"actions — goes to your human via progress while you keep waiting (a timeout_at "
|
||||
"actions — is escalated with the escalate verb (parent-first; the reply "
|
||||
"reaches your mailbox on a later round) while you keep waiting (a timeout_at "
|
||||
"question benign-declines at the engine timeout; timeout_at=null waits until answered). "
|
||||
"If your task text instructs you to execute the work natively yourself, that "
|
||||
"instruction described the metered fallback and is superseded by this dispatch. "
|
||||
|
|
|
|||
|
|
@ -51,7 +51,7 @@ from ouroboros.skill_publish_result import apply_skill_publish_receipt_veto
|
|||
from ouroboros.task_finalization import (
|
||||
build_sealed_final_package,
|
||||
build_swarm_efficiency as _build_swarm_efficiency, # moved (module ceiling); tests import it here
|
||||
deliver_final_message_live, prepare_terminal_send_event, register_final_answer_owed,
|
||||
deliver_final_message_live, prepare_terminal_send_event, register_final_answer_owed, stamp_root_final_phase,
|
||||
sealed_final_prompt_section, terminal_result_fields, # noqa: F401 -- the pipeline module keeps its historical import surface for the synthesis leaf
|
||||
)
|
||||
from ouroboros.dialogue_provenance import is_presence_task, presence_provenance_fields # noqa: F401 -- the pipeline module keeps its historical import surface for the synthesis leaf
|
||||
|
|
@ -528,13 +528,10 @@ def emit_task_results(
|
|||
# used to buffer the send with no delivery_id and no owed registration
|
||||
# at all. Seam + dedup: ouroboros/task_finalization.py.
|
||||
if _is_root_post_task(task) and not _presence:
|
||||
if not _root_post_task_already_completed(env, task):
|
||||
# The owner's answer leaves BEFORE post-task synthesis: the
|
||||
# typed phase marker rides the frame (progress_meta merges
|
||||
# into the WS chat payload) so the client holds the card on
|
||||
# "Finalizing…" until the settled task_done, instead of
|
||||
# reading the early final as the task's terminal conclusion.
|
||||
send_event.setdefault("progress_meta", {})["task_phase"] = "finalizing"
|
||||
stamp_root_final_phase(
|
||||
send_event, task,
|
||||
post_task_open=not _root_post_task_already_completed(env, task),
|
||||
)
|
||||
register_final_answer_owed(task, send_event, env_drive_root=env.drive_root)
|
||||
_store_task_result(
|
||||
env, task, text, usage, llm_trace, review_evidence=review_evidence,
|
||||
|
|
|
|||
|
|
@ -525,14 +525,26 @@ def record_token_density(
|
|||
prompt_tokens: Any,
|
||||
source: str = "dispatch_usage",
|
||||
route_fp: str = "",
|
||||
basis: str = "raw",
|
||||
) -> None:
|
||||
"""Persist one timestamped raw witness, best-effort and write-throttled."""
|
||||
"""Persist one timestamped raw witness, best-effort and write-throttled.
|
||||
|
||||
``basis`` names how ``prompt_chars`` measured image blocks — "raw"
|
||||
(base64 bytes, pre-fix rows) vs "bounded_proxy" (the provider-billing
|
||||
proxy the fit estimator measures on). The row carries it so the two
|
||||
bases can never be silently mixed again by a later "unification".
|
||||
"""
|
||||
fp = str(fingerprint or "").strip()
|
||||
route = str(route_fp or "").strip()
|
||||
density = _density_of(prompt_chars, prompt_tokens)
|
||||
if not fp or density <= 0:
|
||||
return
|
||||
memo_key = f"{fp}\0{route}"
|
||||
# ``basis`` is part of the witness identity end-to-end: the main resolver
|
||||
# accepts only bounded_proxy rows, so a fresh RAW row at the same numeric
|
||||
# density must not throttle the FIRST bounded witness as "no drift" —
|
||||
# that left the resolver cold for the whole freshness window on an
|
||||
# upgraded store (final-lane finding, probe-reproduced).
|
||||
memo_key = f"{fp}\0{route}\0{basis}"
|
||||
memo = _DENSITY_MEMO.get(memo_key)
|
||||
if (
|
||||
memo and _age_seconds(memo[1]) < _TOKEN_DENSITY_FRESH_SEC
|
||||
|
|
@ -550,7 +562,11 @@ def record_token_density(
|
|||
and _age_seconds(str(pair.get("observed_at") or "")) < _TOKEN_DENSITY_TTL_SEC
|
||||
and _density_of(pair.get("prompt_chars"), pair.get("prompt_tokens")) > 0
|
||||
]
|
||||
route_pairs = [pair for pair in pairs if str(pair.get("route_fp") or "") == route]
|
||||
route_pairs = [
|
||||
pair for pair in pairs
|
||||
if str(pair.get("route_fp") or "") == route
|
||||
and str(pair.get("basis") or "raw") == str(basis or "raw")
|
||||
]
|
||||
newest = max(
|
||||
enumerate(route_pairs),
|
||||
key=lambda item: (*_density_recency_key(item[1]), item[0]),
|
||||
|
|
@ -577,6 +593,7 @@ def record_token_density(
|
|||
"observation_seq": observation_seq,
|
||||
"source": str(source or "dispatch_usage"),
|
||||
"route_fp": route,
|
||||
"basis": str(basis or "raw"),
|
||||
})
|
||||
indexed_pairs = list(enumerate(pairs))
|
||||
densest_index, densest = max(
|
||||
|
|
@ -628,13 +645,24 @@ def _normalized_density_model(model_id: str) -> str:
|
|||
return str(model_id or "").strip()
|
||||
|
||||
|
||||
def _fresh_density_pairs(store: Dict[str, Any], model_id: str = "") -> List[Tuple[Dict[str, Any], float]]:
|
||||
def _fresh_density_pairs(
|
||||
store: Dict[str, Any], model_id: str = "", *, basis: str = "",
|
||||
) -> List[Tuple[Dict[str, Any], float]]:
|
||||
# ``basis`` filters to rows measured on one named basis. The MAIN fit
|
||||
# resolver passes "bounded_proxy" — its multiplier must match the fit
|
||||
# estimator's own measure: a pre-basis row (no stamp) or a legacy ``raw``
|
||||
# row was measured against raw base64 chars and can sit at 0.05-0.65 on
|
||||
# image routes, so letting it stay authoritative for its 14-day TTL after
|
||||
# an upgrade re-poisons exactly what the basis fix cures (the cost is a
|
||||
# brief cold start at 1.0). Review/aggregate resolvers pass no basis: their
|
||||
# text-heavy witnesses measure the same on either basis.
|
||||
entries = [store.get(model_id) or {}] if model_id else list(store.values())
|
||||
return [
|
||||
(pair, density)
|
||||
for entry in entries if isinstance(entry, dict)
|
||||
for pair in (entry.get("pairs") or []) if isinstance(pair, dict)
|
||||
if _age_seconds(str(pair.get("observed_at") or "")) < _TOKEN_DENSITY_TTL_SEC
|
||||
and (not basis or str(pair.get("basis") or "") == basis)
|
||||
for density in [_density_of(pair.get("prompt_chars"), pair.get("prompt_tokens"))]
|
||||
if density > 0
|
||||
]
|
||||
|
|
@ -646,7 +674,7 @@ def resolve_main_token_density(drive_root: Any, route_fp: str, model_id: str) ->
|
|||
store = _load(drive_root).get("token_density", {}) or {}
|
||||
route = str(route_fp or "").strip()
|
||||
route_pairs = [
|
||||
item for item in _fresh_density_pairs(store)
|
||||
item for item in _fresh_density_pairs(store, basis="bounded_proxy")
|
||||
if route and str(item[0].get("route_fp") or "") == route
|
||||
]
|
||||
if route_pairs:
|
||||
|
|
@ -654,7 +682,9 @@ def resolve_main_token_density(drive_root: Any, route_fp: str, model_id: str) ->
|
|||
enumerate(route_pairs),
|
||||
key=lambda item: (*_density_recency_key(item[1][0]), item[0]),
|
||||
)[1][1], "fresh_route_usage"
|
||||
model_pairs = _fresh_density_pairs(store, _normalized_density_model(model_id))
|
||||
model_pairs = _fresh_density_pairs(
|
||||
store, _normalized_density_model(model_id), basis="bounded_proxy",
|
||||
)
|
||||
if model_pairs:
|
||||
return max(
|
||||
enumerate(model_pairs),
|
||||
|
|
@ -999,3 +1029,50 @@ def probe(
|
|||
detail="no provider metadata; owner-ack required for a >=1M gate")
|
||||
_store_evidence(drive_root, "probes", fp, ev.to_json())
|
||||
return ev
|
||||
|
||||
|
||||
# Cache-inclusive prompt totals are measurable; GigaChat's semantics remain unknown.
|
||||
_CACHE_INCLUSIVE_PROMPT_TOKEN_PROVIDERS = frozenset({
|
||||
"openrouter", "openai", "openai-compatible", "cloudru", "local", "anthropic",
|
||||
})
|
||||
|
||||
|
||||
def observe_token_density(request: Any, usage: Optional[Dict[str, Any]], *, drive_root_resolver: Any) -> None:
|
||||
"""Learn density after settlement; unknown cache semantics produce no witness."""
|
||||
try:
|
||||
normalized = dict(usage or {})
|
||||
cache_bearing = bool(
|
||||
int(normalized.get("cached_tokens") or 0)
|
||||
or int(normalized.get("cache_write_tokens") or 0)
|
||||
)
|
||||
provider = str(request.provider or "").strip().lower()
|
||||
if cache_bearing and provider not in _CACHE_INCLUSIVE_PROMPT_TOKEN_PROVIDERS:
|
||||
return
|
||||
real = int(normalized.get("prompt_tokens") or normalized.get("input_tokens") or 0)
|
||||
# The witness MUST calibrate the basis the fit estimator measures on
|
||||
# (bounded image proxy) — the raw-base64 basis fed a self-consistent
|
||||
# ~27% under-prediction: measure_main_fit multiplied a BOUNDED
|
||||
# estimate by a RAW-basis density. `prompt_tokens_estimate` stays raw
|
||||
# on purpose — `_reservation_cost` reads it and budget reservation
|
||||
# wants the conservative over-count (owner decision 3=A); do NOT
|
||||
# unify the two consumers onto one basis.
|
||||
estimate = int(request.prompt_tokens_bounded_estimate or 0)
|
||||
basis = "bounded_proxy"
|
||||
if estimate <= 0:
|
||||
estimate = int(request.prompt_tokens_estimate or 0)
|
||||
basis = "raw"
|
||||
if real <= 0 or estimate <= 0:
|
||||
return
|
||||
from ouroboros.provider_models import normalize_model_identity
|
||||
|
||||
record_token_density(
|
||||
drive_root_resolver(request.drive_root),
|
||||
normalize_model_identity(str(request.model or "")),
|
||||
prompt_chars=estimate * 4,
|
||||
prompt_tokens=real,
|
||||
source="dispatch_usage",
|
||||
route_fp=str(request.physical_context.route_fp if request.physical_context else ""),
|
||||
basis=basis,
|
||||
)
|
||||
except Exception:
|
||||
log.debug("token-density observation skipped", exc_info=True)
|
||||
|
|
|
|||
|
|
@ -1,12 +1,12 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"release": {
|
||||
"version": "3.9.2",
|
||||
"build_sha": "5a5c3eba28b436ed5a5e6ec861e8b8f232bb82d2",
|
||||
"version": "3.9.4",
|
||||
"build_sha": "7402e18cbbad61ca0b4211e3730a5090f3b3fd89",
|
||||
"protocol_major": 3,
|
||||
"archive_url": "https://github.com/razzant/claudexor/releases/download/v3.9.2/claudexor-runtime-3.9.2.tar.gz",
|
||||
"sha256": "c56cf6c44ca06ea527f687d98a1759ddfc7fdd6600d5d7232dbb26004763b8dc",
|
||||
"size_bytes": 21975797,
|
||||
"archive_url": "https://github.com/razzant/claudexor/releases/download/v3.9.4/claudexor-runtime-3.9.4.tar.gz",
|
||||
"sha256": "1f89db6920679b5bd679eca4be4667910aa2c63ac8444c7c1578d2df088872cf",
|
||||
"size_bytes": 21998999,
|
||||
"node_version": "24.16.0",
|
||||
"node_artifacts": {
|
||||
"darwin-arm64": {
|
||||
|
|
|
|||
|
|
@ -174,6 +174,8 @@ CLAUDEXOR_MIN_VERSION: str = "3.2.0"
|
|||
# not buy a lane-wide refusal (AGENTS.md "Disclose instead of forbid"). Two gates, two
|
||||
# questions, no overlap; bands: docs/DELEGATED_ADMISSION.md.
|
||||
CLAUDEXOR_DELEGATED_MARKER_MIN_VERSION: str = "3.3.0"
|
||||
# Engine floor for the delegated ``workspaceRoot`` field (#362 stable-target routes).
|
||||
CLAUDEXOR_DELEGATED_WORKSPACE_ROOT_MIN_VERSION: str = "3.8.1"
|
||||
|
||||
|
||||
# Boot-time runtime-mode baseline. Pinning the owner-selected mode after settings load stops an
|
||||
|
|
|
|||
|
|
@ -1273,11 +1273,15 @@ class BackgroundConsciousness:
|
|||
self._registry._ctx.pending_events = []
|
||||
self._registry._ctx.event_queue = self._event_queue
|
||||
self._registry._ctx.task_id = "bg-consciousness"
|
||||
from ouroboros.tool_capabilities import BACKGROUND_DELEGATION_ROLE
|
||||
|
||||
self._registry._ctx.task_metadata = {
|
||||
"root_task_id": "bg-consciousness",
|
||||
"session_id": "background-consciousness",
|
||||
"actor_id": "background-consciousness",
|
||||
"delegation_role": "background",
|
||||
# Owner-delivery gating keys on this role: BG frames must stay
|
||||
# cycle-end deferred (pause discipline), never live mid-cycle.
|
||||
"delegation_role": BACKGROUND_DELEGATION_ROLE,
|
||||
}
|
||||
|
||||
timeout_sec = self._registry.get_timeout(fn_name)
|
||||
|
|
|
|||
|
|
@ -1219,7 +1219,9 @@ def _capture_context_core(
|
|||
|
||||
semi_stable_text = "\n\n".join(semi_stable_parts)
|
||||
|
||||
health_section = build_health_invariants(context_env)
|
||||
health_section = build_health_invariants(
|
||||
context_env, task_id=str(task.get("id") or "")
|
||||
)
|
||||
dynamic_parts = []
|
||||
if health_section:
|
||||
dynamic_parts.append(health_section)
|
||||
|
|
|
|||
|
|
@ -250,3 +250,25 @@ BG_OBSERVATIONS_WARN_BYTES = 20_000_000
|
|||
# explicit full-history read becomes seconds-scale; this is observability, not
|
||||
# a retention gate and never shortens the memory horizon.
|
||||
CHAT_ARCHIVE_SCAN_WARN_BYTES = 100_000_000
|
||||
|
||||
|
||||
def estimate_message_chars(messages: Any) -> int:
|
||||
"""Message chars with image blocks at the provider-billing proxy.
|
||||
|
||||
The bounded basis shared by the fit estimator, the density witness and
|
||||
the compaction proxy — image base64 never counts as text here.
|
||||
"""
|
||||
total = 0
|
||||
for msg in messages:
|
||||
content = msg.get("content")
|
||||
if isinstance(content, list):
|
||||
for block in content:
|
||||
if not isinstance(block, dict):
|
||||
continue
|
||||
if str(block.get("type") or "") in ("image_url", "image"):
|
||||
total += IMAGE_BLOCK_CHAR_EQUIVALENT
|
||||
continue
|
||||
total += len(str(block.get("text", "")))
|
||||
else:
|
||||
total += len(str(content or ""))
|
||||
return total
|
||||
|
|
|
|||
|
|
@ -212,6 +212,23 @@ def tool_schema_tokens(tools: Optional[List[Dict[str, Any]]]) -> int:
|
|||
return estimate_tokens(json.dumps(tools, ensure_ascii=False, sort_keys=True, default=str))
|
||||
|
||||
|
||||
def bounded_prompt_tokens_for_payload(prompt_payload: Dict[str, Any], fallback_chars: int) -> int:
|
||||
"""The density witness's basis: the fit estimator's own token count for a
|
||||
request payload (messages + tools, images at the proxy), or ``fallback_chars
|
||||
// 4`` when there is no message list. Kept beside the estimator so the two
|
||||
can never diverge; ``estimate_message_chars`` dropped tool_call objects and
|
||||
made density ~1.4x high on the tool-heavy shape (measure_main_fit multiplies
|
||||
THIS quantity)."""
|
||||
try:
|
||||
messages = prompt_payload.get("messages")
|
||||
if isinstance(messages, list):
|
||||
return int(estimate_context_prompt_tokens(
|
||||
messages, prompt_payload.get("tools") or prompt_payload.get("functions")))
|
||||
except Exception:
|
||||
pass
|
||||
return max(0, int(fallback_chars) // 4)
|
||||
|
||||
|
||||
def estimate_context_prompt_tokens(
|
||||
messages: List[Dict[str, Any]],
|
||||
tools: Optional[List[Dict[str, Any]]] = None,
|
||||
|
|
|
|||
|
|
@ -193,7 +193,18 @@ def _stray_server_note(env: Any) -> str:
|
|||
return note
|
||||
|
||||
|
||||
def build_health_invariants(env: Any) -> str:
|
||||
def build_health_invariants(env: Any, task_id: str = "") -> str:
|
||||
"""Render the health-invariant WARNING block for one reader's context.
|
||||
|
||||
``task_id`` names the READING task so delegated-run obligations can shape
|
||||
their instruction clause by ownership: the obligations stay globally
|
||||
visible (owner doctrine — a preserved-and-invisible result is how work
|
||||
rots on disk), but only the OWNER task receives the call-shaped
|
||||
instruction. A non-owner told to call ``integrate_delegated_patch`` gets
|
||||
a structural ``run_not_owned`` refusal and an obligation it can never
|
||||
discharge. Empty ``task_id`` (Background Consciousness, legacy callers)
|
||||
keeps the call-shaped wording — an unattributed reader may be the owner.
|
||||
"""
|
||||
import time as _time
|
||||
|
||||
checks: List[str] = []
|
||||
|
|
@ -339,12 +350,25 @@ def build_health_invariants(env: Any) -> str:
|
|||
# nothing is live and nothing is mutating — but it stays visible until the
|
||||
# read happens, which is the whole difference between a disclosure and a fact
|
||||
# someone acts on. It clears itself the moment the acknowledgement lands.
|
||||
# The instruction clause is ownership-aware: the staged artifact lives
|
||||
# under the OWNER's task drive and the read acknowledgement only
|
||||
# credits the owner, so telling a foreign task to read it mints an
|
||||
# obligation it structurally cannot discharge.
|
||||
if task_id and run.task_id and task_id != run.task_id:
|
||||
action = (
|
||||
f"Only its owner task {run.task_id} can read and acknowledge it; "
|
||||
f"note the gap honestly instead of attempting the read here."
|
||||
)
|
||||
else:
|
||||
action = (
|
||||
"Read it with read_file root='task_drive' until the artifact is "
|
||||
"covered end to end, or say plainly that the result was not "
|
||||
"collected."
|
||||
)
|
||||
checks.append(
|
||||
f"WARNING: DELEGATED RESULT NEVER READ — run {run.run_id or '?'} settled "
|
||||
f"with its full output staged at {run.output_artifact} and never read to "
|
||||
f"EOF (owner task {run.task_id or '?'}). Read it with read_file "
|
||||
f"root='task_drive' until the artifact is covered end to end, or say "
|
||||
f"plainly that the result was not collected."
|
||||
f"EOF (owner task {run.task_id or '?'}). {action}"
|
||||
)
|
||||
except Exception:
|
||||
pass
|
||||
|
|
@ -376,11 +400,24 @@ def build_health_invariants(env: Any) -> str:
|
|||
f"integrate_delegated_patch captures it at disposition)"
|
||||
)
|
||||
persist_clause = "The snapshot persists until that disposition."
|
||||
# Ownership-aware instruction: integrate_delegated_patch refuses a
|
||||
# non-owner with run_not_owned and only the owner can write the
|
||||
# clearing PATCH_DISPOSED row, so a foreign reader is told WHO must
|
||||
# act instead of being handed a call that structurally refuses.
|
||||
if task_id and run.task_id and task_id != run.task_id:
|
||||
decide_clause = (
|
||||
f"Only its owner task {run.task_id} can decide it (a foreign "
|
||||
f"integrate_delegated_patch call is refused as run_not_owned)."
|
||||
)
|
||||
else:
|
||||
decide_clause = (
|
||||
f"decide with integrate_delegated_patch(run_id='{run.run_id}', "
|
||||
f"decision='apply'|'reject')."
|
||||
)
|
||||
checks.append(
|
||||
f"WARNING: DELEGATED PATCH AWAITS DISPOSITION — run {run.run_id or '?'} "
|
||||
f"(owner task {run.task_id or '?'}) {state_clause} and no apply/reject "
|
||||
f"recorded. Nothing reaches the shared tree by itself: decide with "
|
||||
f"integrate_delegated_patch(run_id='{run.run_id}', decision='apply'|'reject'). "
|
||||
f"recorded. Nothing reaches the shared tree by itself: {decide_clause} "
|
||||
f"{persist_clause}"
|
||||
)
|
||||
except Exception:
|
||||
|
|
|
|||
|
|
@ -311,3 +311,18 @@ def review_transport_timeout(model: Any, explicit: Any = None, deadline_at: Any
|
|||
deadline_at=deadline or None,
|
||||
reserve_sec=get_finalization_grace_sec(),
|
||||
)
|
||||
|
||||
|
||||
def deadline_expired(ctx) -> bool:
|
||||
"""True when the task HAS a deadline and it has already passed.
|
||||
|
||||
The distinction ``deadline_remaining_sec`` alone cannot make: it answers
|
||||
0.0 both for "no deadline" and for "the deadline is behind us", and
|
||||
collapsing them let an EXPIRED nanny hand a fresh run the absolute task
|
||||
ceiling.
|
||||
"""
|
||||
meta = getattr(ctx, "task_metadata", {})
|
||||
meta = meta if isinstance(meta, dict) else {}
|
||||
if parse_deadline_ts(meta.get("deadline_at")) is None:
|
||||
return False
|
||||
return deadline_remaining_sec(ctx) <= 0
|
||||
|
|
|
|||
|
|
@ -121,6 +121,8 @@ class RunCustody:
|
|||
profile_id: str = ""
|
||||
project_id: str = ""
|
||||
project_owned: bool = False
|
||||
# #362: a stable user-target registration outlives any single run.
|
||||
project_persistent: bool = False
|
||||
root_task_id: str = ""
|
||||
parent_task_id: str = ""
|
||||
category: str = "subagent"
|
||||
|
|
@ -144,6 +146,7 @@ class RunCustody:
|
|||
authority_fingerprint: str = ""
|
||||
ledger_recorded: bool = False
|
||||
settled: bool = False
|
||||
terminal_state: str = "" # SETTLED row's state, replayed (empty pre-existing/CLOSED_ABSENT)
|
||||
containment_disclosed: bool = False # written once; a re-poll must not re-find
|
||||
unread_disclosed: bool = False # settled-never-read omission named durably
|
||||
# Staged-output half of the terminal story (D7). ``output_artifact``:
|
||||
|
|
@ -309,35 +312,12 @@ def _iter_rows(path: pathlib.Path, tail_bytes: Optional[int] = None) -> Iterator
|
|||
return
|
||||
|
||||
|
||||
# The STARTED row's string facts as ``(RunCustody attribute, row key)`` pairs —
|
||||
# one table shared by the replay and the ``record_started`` emit.
|
||||
_STARTED_STR_FIELDS: Tuple[Tuple[str, str], ...] = tuple(
|
||||
(attr, "route" if attr == "route_id" else attr) for attr in (
|
||||
"task_id", "route_id", "model", "profile_id", "project_id", "root_task_id",
|
||||
"parent_task_id", "category", "source", *REVIEW_ATTRIBUTION_KEYS,
|
||||
"ledger_root", "idempotency_key", "invocation_id",
|
||||
"snapshot_id", "execution_root", "baseline_sha", "target_root",
|
||||
"authority_source", "access", "mode", "isolation",
|
||||
"selected_subagent_id", "config_fingerprint", "work_order_fingerprint",
|
||||
"work_order_coverage", "authority_fingerprint",
|
||||
)
|
||||
from ouroboros.delegate_registration_policy import (
|
||||
record_persistent as _record_persistent,
|
||||
STARTED_FIRST_WINS_FACTS as _STARTED_FIRST_WINS_FACTS,
|
||||
STARTED_PROGRESS_FLAGS as _STARTED_PROGRESS_FLAGS,
|
||||
STARTED_STR_FIELDS as _STARTED_STR_FIELDS,
|
||||
)
|
||||
# Progress carried forward from a previous row: an idempotent re-start writes a
|
||||
# SECOND started row; replacing wholesale would forget a settlement and put a
|
||||
# finished run back into the orphan sweep (which would cancel it).
|
||||
_STARTED_PROGRESS_FLAGS: Tuple[str, ...] = (
|
||||
"ledger_recorded", "settled", "containment_disclosed", "unread_disclosed",
|
||||
"output_artifact", "output_complete", "output_sha", "output_consumed",
|
||||
"patch_captured", "patch_disposed", "patch_apply_pending")
|
||||
# Binding/authority facts are FIRST-WINS (R1-2): a later idempotent STARTED row
|
||||
# may be minted by a context that no longer knows the original binding; the
|
||||
# first recorded fact is authoritative and is never erased or retargeted.
|
||||
_STARTED_FIRST_WINS_FACTS: Tuple[str, ...] = (
|
||||
"snapshot_id", "execution_root", "baseline_sha", "target_root",
|
||||
"authority_source", "resource_ref", "selected_subagent_id",
|
||||
"config_fingerprint", "work_order_fingerprint", "work_order_coverage",
|
||||
"authority_fingerprint", "work_order_source_request", "category", "source",
|
||||
*REVIEW_ATTRIBUTION_KEYS)
|
||||
|
||||
from ouroboros.delegate_source_coverage import (
|
||||
apply_source_delivery_confirmation,
|
||||
|
|
@ -361,6 +341,7 @@ def _merge_started_into(entry: RunCustody, previous: RunCustody) -> None:
|
|||
for attr in _STARTED_PROGRESS_FLAGS:
|
||||
setattr(entry, attr, getattr(previous, attr))
|
||||
entry.project_owned = previous.project_owned and entry.project_owned
|
||||
entry.project_persistent = previous.project_persistent or entry.project_persistent
|
||||
for attr in _STARTED_FIRST_WINS_FACTS:
|
||||
prior = getattr(previous, attr)
|
||||
if prior:
|
||||
|
|
@ -387,6 +368,7 @@ def _apply(state: Dict[str, RunCustody], row: Dict[str, Any]) -> None:
|
|||
entry = RunCustody(
|
||||
run_id=run_id,
|
||||
project_owned=bool(row.get("project_owned")),
|
||||
project_persistent=bool(row.get("project_persistent")),
|
||||
delegated=row.get("delegated") is True,
|
||||
resource_ref=dict(ref) if isinstance(ref, dict) else {},
|
||||
work_order_source_request=(
|
||||
|
|
@ -460,14 +442,13 @@ def _apply(state: Dict[str, RunCustody], row: Dict[str, Any]) -> None:
|
|||
disposition = str(row.get("disposition") or "")
|
||||
if disposition:
|
||||
custody.patch_disposed = disposition
|
||||
# A recorded disposition completes the apply-intent story too.
|
||||
custody.patch_apply_pending = False
|
||||
custody.patch_apply_pending = False # a recorded disposition completes the apply-intent story
|
||||
elif kind == SETTLED:
|
||||
# A RUN-level fact: it no longer clears the registration obligation -
|
||||
# PROJECT_RETIRED is the only discharging row (historical logs always
|
||||
# emitted it before SETTLED, so replay is unaffected).
|
||||
custody.ledger_recorded = True
|
||||
custody.settled = True
|
||||
custody.ledger_recorded = custody.settled = True
|
||||
custody.terminal_state = str(row.get("state") or "") or custody.terminal_state
|
||||
elif kind == CLOSED_ABSENT:
|
||||
# Closed, not settled: custody is over, the run leaves ``open_runs``.
|
||||
# The registration survives independently (wholesale clearing here was
|
||||
|
|
@ -692,6 +673,9 @@ def invocation_record(drive_root: Any, invocation_id: str) -> Optional[Dict[str,
|
|||
"route": str(row.get("route") or ""),
|
||||
"project_id": str(row.get("project_id") or ""),
|
||||
"project_owned": bool(row.get("project_owned")),
|
||||
# Absence is a fact: legacy rows fall back to the stored request.
|
||||
**({"project_persistent": bool(row["project_persistent"])}
|
||||
if "project_persistent" in row else {}),
|
||||
"idempotency_key": str(row.get("idempotency_key") or ""),
|
||||
"root_task_id": str(row.get("root_task_id") or ""),
|
||||
"parent_task_id": str(row.get("parent_task_id") or ""),
|
||||
|
|
@ -762,7 +746,7 @@ def record_started(drive_root: Any, custody: RunCustody,
|
|||
# recorded separately can lose half of itself to a crash); shape spreads LAST.
|
||||
return emit(drive_root, STARTED, {
|
||||
"run_id": custody.run_id,
|
||||
"project_owned": custody.project_owned,
|
||||
"project_owned": custody.project_owned, "project_persistent": custody.project_persistent,
|
||||
"resource_ref": custody.resource_ref or {},
|
||||
"work_order_source_request": custody.work_order_source_request or {},
|
||||
**{key: getattr(custody, attr) for attr, key in _STARTED_STR_FIELDS},
|
||||
|
|
@ -800,6 +784,13 @@ def retire_project(drive_root: Any, gateway: Any, custody: RunCustody) -> None:
|
|||
LOWEST-run_id sharer keeps attempting; the rest defer quietly and
|
||||
discharge on the daemon's 404 (deterministic tie-break: someone always
|
||||
attempts)."""
|
||||
if custody.project_persistent:
|
||||
# #362: stable identity outlives the run — discharge the duty DURABLY
|
||||
# (replay must not resurrect owned=True), keep the project itself.
|
||||
custody.project_owned = False
|
||||
emit(drive_root, PROJECT_RETIRED, {"run_id": custody.run_id, "task_id": custody.task_id,
|
||||
"project_id": custody.project_id, "project_kept": True})
|
||||
return
|
||||
if not (custody.project_owned and custody.project_id):
|
||||
return
|
||||
try:
|
||||
|
|
@ -820,6 +811,13 @@ def retire_project(drive_root: Any, gateway: Any, custody: RunCustody) -> None:
|
|||
return
|
||||
rows = [run for run in state.values()
|
||||
if run.project_id == custody.project_id and run.run_id]
|
||||
if any(run.project_persistent for run in rows):
|
||||
# #362: ANY persistent sharer makes the project a durable user
|
||||
# identity — a non-persistent creator must not delete it either.
|
||||
custody.project_owned = False
|
||||
emit(drive_root, PROJECT_RETIRED, {"run_id": custody.run_id, "task_id": custody.task_id,
|
||||
"project_id": custody.project_id, "project_kept": True})
|
||||
return
|
||||
if any(not run.settled and run.run_id != custody.run_id for run in rows):
|
||||
return
|
||||
sharers = sorted(run.run_id for run in rows if run.project_owned)
|
||||
|
|
@ -877,7 +875,8 @@ def settle_run(drive_root: Any, gateway: Any, custody: RunCustody, detail: Dict[
|
|||
settlement is retried; ``settled`` means the durable fact exists."""
|
||||
if custody.settled:
|
||||
return {"settled": True, "ledger_recorded": True,
|
||||
"project_retired": not custody.project_owned, "retried": False}
|
||||
"project_retired": not custody.project_owned and not custody.project_persistent,
|
||||
"project_persistent": custody.project_persistent, "retried": False}
|
||||
summary = summary_of(detail)
|
||||
# Claudexor reports CASH in `spendUsd`, EXACTNESS in `spendEstimated`. A run
|
||||
# is only free when the amount is really zero AND really settled: expired
|
||||
|
|
@ -963,7 +962,8 @@ def settle_run(drive_root: Any, gateway: Any, custody: RunCustody, detail: Dict[
|
|||
return {
|
||||
"settled": custody.settled,
|
||||
"ledger_recorded": custody.ledger_recorded,
|
||||
"project_retired": not custody.project_owned,
|
||||
"project_retired": not custody.project_owned and not custody.project_persistent,
|
||||
"project_persistent": custody.project_persistent,
|
||||
"retried": True,
|
||||
}
|
||||
|
||||
|
|
@ -1057,7 +1057,7 @@ def settled_unread_outputs(drive_root: Any) -> List[RunCustody]:
|
|||
if settled_output_unread(custody)]
|
||||
|
||||
|
||||
def undisposed_patches(drive_root: Any) -> List[RunCustody]:
|
||||
def undisposed_patches(drive_root: Any, state: Optional[Dict[str, RunCustody]] = None) -> List[RunCustody]:
|
||||
"""Settled mutating runs whose snapshot work awaits an explicit apply/reject.
|
||||
|
||||
The C1 counterpart of ``settled_unread_outputs``: a run that executed in a
|
||||
|
|
@ -1069,7 +1069,7 @@ def undisposed_patches(drive_root: Any) -> List[RunCustody]:
|
|||
forever, preserved but findable by nobody. Self-clearing: the
|
||||
``PATCH_DISPOSED`` row flips ``patch_disposed`` in the very replay this reads.
|
||||
"""
|
||||
return [custody for custody in replay(drive_root).values()
|
||||
return [custody for custody in (state if state is not None else replay(drive_root)).values()
|
||||
if custody.snapshot_id and custody.settled and not custody.patch_disposed]
|
||||
|
||||
|
||||
|
|
@ -1207,13 +1207,11 @@ def _cancel_result(drive_root: Any, custody: RunCustody, outcome: str, *, accept
|
|||
# -- reconciliation ------------------------------------------------------------
|
||||
|
||||
|
||||
def owned_project_registrations(drive_root: Any) -> List[RunCustody]:
|
||||
def owned_project_registrations(drive_root: Any, state: Optional[Dict[str, RunCustody]] = None) -> List[RunCustody]:
|
||||
"""Runs whose registration is still owned - settled or not (``open_runs``
|
||||
cannot see registrations that outlive their runs)."""
|
||||
return [
|
||||
custody for custody in replay(drive_root).values()
|
||||
if custody.project_owned and custody.project_id
|
||||
]
|
||||
return [custody for custody in (state if state is not None else replay(drive_root)).values()
|
||||
if custody.project_owned and custody.project_id]
|
||||
|
||||
|
||||
def retire_settled_registrations(drive_root: Any, gateway: Any) -> None:
|
||||
|
|
|
|||
|
|
@ -13,6 +13,7 @@ import logging
|
|||
from typing import Any, Callable, Dict, List, Optional
|
||||
|
||||
from ouroboros._usage_rows import REVIEW_ATTRIBUTION_KEYS
|
||||
from ouroboros.delegate_registration_policy import record_persistent as _record_persistent
|
||||
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
|
|
@ -38,9 +39,11 @@ def _custody():
|
|||
return delegate_custody
|
||||
|
||||
|
||||
def open_runs(drive_root: Any) -> List[RunCustody]:
|
||||
"""Runs with a durable start and no durable settlement."""
|
||||
return [custody for custody in _custody().replay(drive_root).values() if not custody.settled]
|
||||
def open_runs(drive_root: Any, state: Optional[Dict[str, "RunCustody"]] = None) -> List[RunCustody]:
|
||||
"""Runs with a durable start and no durable settlement (``state``: a shared
|
||||
pre-replayed snapshot, so a batch of audits pays one log traversal)."""
|
||||
return [custody for custody in (state if state is not None else _custody().replay(drive_root)).values()
|
||||
if not custody.settled]
|
||||
|
||||
|
||||
def pending_invocations(drive_root: Any,
|
||||
|
|
@ -238,6 +241,7 @@ def _recover_pending_invocation(drive_root: Any, gateway: Any,
|
|||
model=str(body.get("model") or ""),
|
||||
profile_id=str(body.get("credentialProfileId") or ""),
|
||||
project_id=record["project_id"], project_owned=bool(record["project_owned"]),
|
||||
project_persistent=_record_persistent(record),
|
||||
root_task_id=str(record.get("root_task_id") or ""),
|
||||
parent_task_id=str(record.get("parent_task_id") or ""),
|
||||
category=str(record.get("category") or "subagent"),
|
||||
|
|
@ -286,7 +290,7 @@ def _recover_pending_invocation(drive_root: Any, gateway: Any,
|
|||
|
||||
def _retire_recovered_registration(gateway: Any, record: Dict[str, Any]) -> bool:
|
||||
"""Discharge the registration an ORIGINAL attempt owned, when its invocation dies."""
|
||||
if not (record.get("project_owned") and record.get("project_id")):
|
||||
if _record_persistent(record) or not (record.get("project_owned") and record.get("project_id")):
|
||||
return False
|
||||
try:
|
||||
gateway.remove_project(record["project_id"])
|
||||
|
|
|
|||
|
|
@ -92,6 +92,7 @@ def task_execution_evidence(drive_root: Any, task_id: str) -> Dict[str, Any]:
|
|||
succeeded: set = set()
|
||||
failure_states: List[str] = []
|
||||
models: List[str] = []
|
||||
applied_access_profiles: List[str] = []
|
||||
cost_total, cost_known, cost_estimated = 0.0, True, False
|
||||
# Scope finding (a5e59bdf gate): an UNREADABLE log must not collapse into
|
||||
# the same zero-count result as a proven empty one — a reader would then
|
||||
|
|
@ -164,6 +165,21 @@ def task_execution_evidence(drive_root: Any, task_id: str) -> Dict[str, Any]:
|
|||
model = str(row.get("model") or "")
|
||||
if model and model not in models:
|
||||
models.append(model)
|
||||
# D29 applied-access disclosure: the settlement row records the
|
||||
# ACCESS the engine actually served (effectiveAccess), distinct
|
||||
# from the STARTED row's granted shape. Projected so the chip and
|
||||
# acceptance readers can answer asked-vs-applied without joining
|
||||
# the ledger. Empty when telemetry predates the receipt.
|
||||
# Bounded to the known access-profile vocabulary and a short cap:
|
||||
# the value rides a durable evidence row and the chip tooltip, so a
|
||||
# garbage/oversized engine string must not reach either verbatim.
|
||||
applied_access = str(row.get("access_profile") or "")[:32]
|
||||
if (
|
||||
applied_access
|
||||
and applied_access not in applied_access_profiles
|
||||
and len(applied_access_profiles) < 8
|
||||
):
|
||||
applied_access_profiles.append(applied_access)
|
||||
# A partial work-order run is not a successful delegated substrate until the
|
||||
# durable verified-range union covers the complete canonical brief. Reuse the
|
||||
# custody replay (the same SSOT used by wait/apply) so a restart cannot turn a
|
||||
|
|
@ -214,6 +230,9 @@ def task_execution_evidence(drive_root: Any, task_id: str) -> Dict[str, Any]:
|
|||
# dropped: an estimated sum must never render as an exact receipt.
|
||||
"subscription_cost_estimated": bool(settled and cost_known and cost_estimated),
|
||||
"harness_models": models,
|
||||
# Additive (D29 projection): the access the engine actually served,
|
||||
# from SETTLED rows only. Empty list = no receipt disclosed it.
|
||||
"applied_access_profiles": applied_access_profiles,
|
||||
}
|
||||
|
||||
|
||||
|
|
@ -289,6 +308,67 @@ def acceptance_substrate_facts(ctx: Any, task_id: str) -> Dict[str, Any]:
|
|||
|
||||
|
||||
_ACCEPT_DELTA_CHILD_CAP = 20 # reduced-children rows in the finalizer aggregate
|
||||
_ACCEPT_PATCH_DISPOSITION_CAP = 20 # disposition rows in the acceptance section
|
||||
|
||||
|
||||
def acceptance_patch_dispositions(drive_root: Any, task_id: str) -> Dict[str, Any]:
|
||||
"""Typed aggregate of this parent's patch apply/reject decisions (D-trace).
|
||||
|
||||
``integrate_delegated_patch`` applied a child's diff on FIVE mechanical
|
||||
manifest fields with zero review facts anywhere on the path; the owner
|
||||
decision (4=A) is to ATTEST the apply rather than invent a review: every
|
||||
verdict now also lands as a ``subagent_patch_verdict`` custody row, and
|
||||
this section projects those rows host-attested into the acceptance packet
|
||||
(the ``capability_deltas`` shape — bounded, first-class, never squeezed
|
||||
through the 4KB artifact-preview cliff). ABSENCE of the section means "no
|
||||
disposition recorded", never "reviewed clean"; an unreadable log is the
|
||||
typed ``evidence_read_failed`` marker, never an empty-therefore-clean
|
||||
section (the ``task_execution_evidence`` rule, GR6-4).
|
||||
"""
|
||||
from ouroboros import delegate_custody as custody
|
||||
from ouroboros.utils import truncate_review_artifact
|
||||
|
||||
tid = str(task_id or "")
|
||||
out: Dict[str, Any] = {}
|
||||
log_path = custody.event_log_path(drive_root)
|
||||
try:
|
||||
if log_path.exists():
|
||||
with log_path.open("rb"):
|
||||
pass
|
||||
else:
|
||||
return out
|
||||
except OSError:
|
||||
return {"evidence_read_failed": True}
|
||||
rows: List[Dict[str, Any]] = []
|
||||
for row in custody._iter_rows(log_path):
|
||||
if str(row.get("type") or "") != "delegate_run_patch_verdict":
|
||||
continue
|
||||
if str(row.get("task_id") or "") != tid:
|
||||
continue
|
||||
rows.append({
|
||||
"child": str(row.get("child_task_id") or ""),
|
||||
"pipeline": str(row.get("pipeline") or ""),
|
||||
"disposition": str(row.get("disposition") or ""),
|
||||
"applied": bool(row.get("applied")),
|
||||
"reason": truncate_review_artifact(str(row.get("reason") or ""), limit=600),
|
||||
"patch_sha256": str(row.get("patch_sha256") or ""),
|
||||
**({"verdict_artifact_write_failed": True}
|
||||
if row.get("verdict_artifact_write_failed") else {}),
|
||||
})
|
||||
if not rows:
|
||||
return out
|
||||
out["total"] = len(rows)
|
||||
# The honest headline the panel weighs — computed over the COMPLETE row
|
||||
# set BEFORE bounding: a delegated apply among the omitted-oldest rows is
|
||||
# exactly the fact the owner's attest decision (4=A) exists to surface,
|
||||
# and deriving it from the truncated view false-negatives past the cap.
|
||||
if any(r["applied"] and r.get("pipeline") == "delegated" for r in rows):
|
||||
out["unreviewed_delegated_apply"] = True
|
||||
if len(rows) > _ACCEPT_PATCH_DISPOSITION_CAP:
|
||||
out["omitted"] = len(rows) - _ACCEPT_PATCH_DISPOSITION_CAP
|
||||
rows = rows[-_ACCEPT_PATCH_DISPOSITION_CAP:]
|
||||
out["rows"] = rows
|
||||
return out
|
||||
|
||||
|
||||
def acceptance_capability_deltas(drive_root: Any, task_id: str, root_task_id: str) -> Dict[str, Any]:
|
||||
|
|
|
|||
|
|
@ -31,7 +31,7 @@ log = logging.getLogger(__name__)
|
|||
# `waiting_on_user` payload. Process-local by design: the durable fact is the
|
||||
# engine's own pending-interaction store, and the only cost of losing this memo
|
||||
# (worker restart) is one duplicate immediate return. Without it, a nanny that
|
||||
# deliberately escalated a question to its human and re-waited would busy-loop —
|
||||
# deliberately escalated a question up the task hierarchy and re-waited would busy-loop —
|
||||
# every delegate_wait would return instantly with the same known question.
|
||||
_REPORTED_INTERACTIONS: Dict[str, frozenset] = {}
|
||||
_REPORTED_INTERACTIONS_MAX_KEYS = 128
|
||||
|
|
@ -128,9 +128,12 @@ def _waiting_on_user_note(pending: List[Dict[str, Any]]) -> str:
|
|||
"delegate_answer(run_id, interaction_id, answers=[{question_id, "
|
||||
"selected_labels, free_text}]) — answer from the task context you "
|
||||
"already hold. A question ABOVE your authority (spending money, "
|
||||
"changing scope, external actions) is not yours to guess: surface it "
|
||||
"to your human via a progress message and keep waiting with "
|
||||
"delegate_wait. If the question carries a source-request envelope, answer "
|
||||
"changing scope, external actions) is not yours to guess: escalate it "
|
||||
"with escalate(question, options, stake, assumption) — as a subagent "
|
||||
"you escalate to your PARENT task (the owner sees only what no "
|
||||
"ancestor answers), the reply reaches your mailbox on a later round, "
|
||||
"and you relay it back with delegate_answer — meanwhile keep waiting "
|
||||
"with delegate_wait. If the question carries a source-request envelope, answer "
|
||||
"with source_response={schema:1, kind:'source_response', complete_sha256, "
|
||||
"source, start_char, end_char, text}; the host verifies the exact canonical "
|
||||
"range before delivering it. Do not promote a partial preview to complete "
|
||||
|
|
|
|||
|
|
@ -31,6 +31,10 @@ def pending_invocations(
|
|||
"route": str(row.get("route") or ""),
|
||||
"project_id": str(row.get("project_id") or ""),
|
||||
"project_owned": bool(row.get("project_owned")),
|
||||
# Absence is a fact (legacy rows): the recovery fallback
|
||||
# derives persistence from the stored request.
|
||||
**({"project_persistent": bool(row["project_persistent"])}
|
||||
if "project_persistent" in row else {}),
|
||||
"idempotency_key": str(row.get("idempotency_key") or ""),
|
||||
"root_task_id": str(row.get("root_task_id") or ""),
|
||||
"parent_task_id": str(row.get("parent_task_id") or ""),
|
||||
|
|
|
|||
|
|
@ -27,6 +27,7 @@ log = logging.getLogger(__name__)
|
|||
|
||||
_TIMELINE_TAIL = 12
|
||||
_TIMELINE_LABEL_CHARS = 300
|
||||
_TEXT_KINDS = ("thinking", "message")
|
||||
# The closing `note` is appended AFTER the advance list is sized, so its cost is
|
||||
# reserved rather than measured; it is a fixed string of this module's own authorship.
|
||||
_NOTE_RESERVE_CHARS = 700
|
||||
|
|
@ -63,11 +64,19 @@ def _bounded(rows: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
|||
return truncate_review_artifact(value, _TIMELINE_LABEL_CHARS) \
|
||||
if isinstance(value, str) else value
|
||||
|
||||
return [
|
||||
{"type": _label(row.get("type")), "title": _label(row.get("title")),
|
||||
"severity": _label(row.get("severity"))}
|
||||
for row in rows[-_TIMELINE_TAIL:]
|
||||
]
|
||||
out = []
|
||||
for row in rows[-_TIMELINE_TAIL:]:
|
||||
item = {"type": _label(row.get("type")), "title": _label(row.get("title")),
|
||||
"severity": _label(row.get("severity"))}
|
||||
if row.get("textKind") in _TEXT_KINDS and isinstance(row.get("detail"), str):
|
||||
# The typed body preserves stream whitespace; the engine's title is
|
||||
# only a preview. Keep the same disclosed bound, without a second copy.
|
||||
item.update(title=_label(row["detail"]), textKind=row["textKind"],
|
||||
textDelta=row.get("textDelta") is True,
|
||||
attemptId=_label(row.get("attemptId")),
|
||||
harnessId=_label(row.get("harnessId")))
|
||||
out.append(item)
|
||||
return out
|
||||
|
||||
|
||||
def timeline_tail(detail: Dict[str, Any]) -> List[Dict[str, Any]]:
|
||||
|
|
@ -526,9 +535,26 @@ def _fitted_pending(payload: Dict[str, Any], pending: List[Dict[str, Any]],
|
|||
|
||||
|
||||
def live_line(run_id: str, advance: _Advance) -> str:
|
||||
titles = " · ".join(
|
||||
str(row.get("title") or row.get("type") or "") for row in advance.events
|
||||
)
|
||||
parts: List[str] = []
|
||||
previous_stream = None
|
||||
for row in advance.events:
|
||||
text_kind = row.get("textKind")
|
||||
is_text = text_kind in _TEXT_KINDS and isinstance(row.get("title"), str)
|
||||
text = row["title"] if is_text else str(row.get("title") or row.get("type") or "")
|
||||
stream = (text_kind, row.get("attemptId"), row.get("harnessId")) \
|
||||
if is_text and row.get("textDelta") is True else None
|
||||
if not text:
|
||||
if stream != previous_stream:
|
||||
previous_stream = None
|
||||
continue
|
||||
# Native fragments may split a word or contain only whitespace. Only the
|
||||
# engine's explicit delta fact authorizes concatenation, never the prose.
|
||||
if parts and stream is not None and stream == previous_stream:
|
||||
parts[-1] += text
|
||||
else:
|
||||
parts.append(text)
|
||||
previous_stream = stream
|
||||
titles = " · ".join(parts)
|
||||
# A batch larger than the display tail shows its LAST rows. Saying how many it is not
|
||||
# showing is the difference between an honest subset and a line that reads like the
|
||||
# whole batch — the human's stream is the one surface that had no marker at all. It
|
||||
|
|
@ -576,7 +602,7 @@ def window_payload(
|
|||
|
||||
``pending_interactions`` is the caller's BOUNDED question projection (B4): a
|
||||
known question the model chose not to answer keeps riding every expiry payload
|
||||
beside the ``waiting_on_user`` boolean, so a nanny that escalated to its human
|
||||
beside the ``waiting_on_user`` boolean, so a nanny that escalated up the hierarchy
|
||||
is re-shown what the run is paused on instead of a bare flag. MEASURED into
|
||||
the payload in BOTH branches (F2) — through ``_fitted_pending``'s shed
|
||||
discipline, before the advance list is sized — so it can never push the
|
||||
|
|
@ -602,7 +628,9 @@ def window_payload(
|
|||
payload["note"] = (
|
||||
("The run is alive and PAUSED on the question(s) it already asked "
|
||||
"(waiting_on_user; see pending_interactions). Decide: answer with "
|
||||
"delegate_answer, escalate to your human via a progress message, or "
|
||||
"delegate_answer, escalate an above-authority question with the "
|
||||
"escalate verb (parent-first; the reply reaches your mailbox on a "
|
||||
"later round), or "
|
||||
f"keep waiting (call again) — {waiting_expiry_clause(pending_interactions)}. "
|
||||
"Do not cancel a run merely because it asked a question.")
|
||||
if waiting_on_user else
|
||||
|
|
@ -664,7 +692,8 @@ def window_payload(
|
|||
# The expiry claim is keyed on the rows' own timeout_at (R2-7e).
|
||||
payload["note"] += (
|
||||
" The run is PAUSED on a question: answer it (delegate_answer), "
|
||||
f"escalate to your human, or keep waiting — "
|
||||
"raise it with the escalate verb (parent-first) if it is above "
|
||||
f"your authority, or keep waiting — "
|
||||
f"{waiting_expiry_clause(pending_interactions)}; do not cancel over it."
|
||||
)
|
||||
return payload
|
||||
|
|
|
|||
101
ouroboros/delegate_registration_policy.py
Normal file
101
ouroboros/delegate_registration_policy.py
Normal file
|
|
@ -0,0 +1,101 @@
|
|||
"""Registration persistence policy and the STARTED-row field tables.
|
||||
|
||||
The leaf module the delegate custody core leans on at the module line gate
|
||||
(both ``delegate_custody.py`` and ``tools/delegate.py`` sit exactly on the
|
||||
1600-line ceiling): the #362 persistence decision lives here, and the
|
||||
STARTED-row shape tables move here with it so the core pays for the marker
|
||||
by extraction, not compression.
|
||||
|
||||
#362 (the f9356572 A3 remediation, ported): on engines that support the
|
||||
delegated ``workspaceRoot`` field the run registers the user's STABLE target
|
||||
root as its project — an identity that outlives any individual run. The old
|
||||
retire class deleted that registration at settlement (custody rows,
|
||||
settle_run, pending invocations, retry binding, the orphan sweep), so a
|
||||
user's own project vanished when a delegated run finished. A registration
|
||||
marked persistent survives every retirement path; only ownership of the
|
||||
one-shot snapshot registrations is discharged.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Tuple
|
||||
|
||||
from ouroboros._usage_rows import REVIEW_ATTRIBUTION_KEYS
|
||||
|
||||
|
||||
def persistent_registration(execution_root: str, access: str) -> bool:
|
||||
"""Is this run's project registration a durable user identity?
|
||||
|
||||
True exactly when the engine bound a stable execution workspace
|
||||
(``workspaceRoot`` supported => non-empty execution root) and the run
|
||||
writes into the user's own tree (``workspace_write``): that registration
|
||||
names the user's project, not a disposable snapshot, and must outlive
|
||||
the run (#362).
|
||||
"""
|
||||
return bool(str(execution_root or "").strip()) and str(access or "") == "workspace_write"
|
||||
|
||||
|
||||
def record_persistent(record) -> bool:
|
||||
"""Persistence of a DURABLE record, with a pre-marker upgrade fallback.
|
||||
|
||||
Rows written before the marker existed carry no ``project_persistent``
|
||||
key; deriving from the immutable stored REQUEST (the delegated
|
||||
``execution.workspaceRoot`` plus ``access``) keeps an old pending retry
|
||||
from deleting the user's stable project. Rows that carry the key are
|
||||
authoritative (first-wins doctrine: never recompute from a live engine).
|
||||
"""
|
||||
if not isinstance(record, dict):
|
||||
return False
|
||||
if "project_persistent" in record:
|
||||
return bool(record.get("project_persistent"))
|
||||
request = record.get("request")
|
||||
if not isinstance(request, dict):
|
||||
return False
|
||||
execution = request.get("execution")
|
||||
workspace_root = execution.get("workspaceRoot") if isinstance(execution, dict) else ""
|
||||
return persistent_registration(str(workspace_root or ""), str(request.get("access") or ""))
|
||||
|
||||
|
||||
def resolve_registration(gateway, scope_root: str, execution_root: str, access: str):
|
||||
"""Register (or adopt) the scope root and decide the marker in one step.
|
||||
|
||||
Returns ``(project_id, owned_project_id, project_persistent)``: an
|
||||
existing registration is adopted unowned; a fresh one is owned by this
|
||||
start; #362 marks the stable-target case persistent so no retire path
|
||||
deletes the user's own project.
|
||||
"""
|
||||
existing_project = gateway.find_project_id(scope_root)
|
||||
project_id = existing_project or gateway.register_project(scope_root)
|
||||
owned_project_id = "" if existing_project else project_id
|
||||
return project_id, owned_project_id, persistent_registration(execution_root, access)
|
||||
|
||||
|
||||
# The STARTED row's string facts as ``(RunCustody attribute, row key)`` pairs —
|
||||
# one table shared by the replay and the ``record_started`` emit.
|
||||
STARTED_STR_FIELDS: Tuple[Tuple[str, str], ...] = tuple(
|
||||
(attr, "route" if attr == "route_id" else attr) for attr in (
|
||||
"task_id", "route_id", "model", "profile_id", "project_id", "root_task_id",
|
||||
"parent_task_id", "category", "source", *REVIEW_ATTRIBUTION_KEYS,
|
||||
"ledger_root", "idempotency_key", "invocation_id",
|
||||
"snapshot_id", "execution_root", "baseline_sha", "target_root",
|
||||
"authority_source", "access", "mode", "isolation",
|
||||
"selected_subagent_id", "config_fingerprint", "work_order_fingerprint",
|
||||
"work_order_coverage", "authority_fingerprint",
|
||||
)
|
||||
)
|
||||
# Progress carried forward from a previous row: an idempotent re-start writes a
|
||||
# SECOND started row; replacing wholesale would forget a settlement and put a
|
||||
# finished run back into the orphan sweep (which would cancel it).
|
||||
STARTED_PROGRESS_FLAGS: Tuple[str, ...] = (
|
||||
"ledger_recorded", "settled", "containment_disclosed", "unread_disclosed",
|
||||
"output_artifact", "output_complete", "output_sha", "output_consumed",
|
||||
"patch_captured", "patch_disposed", "patch_apply_pending")
|
||||
# Binding/authority facts are FIRST-WINS (R1-2): a later idempotent STARTED row
|
||||
# may be minted by a context that no longer knows the original binding; the
|
||||
# first recorded fact is authoritative and is never erased or retargeted.
|
||||
STARTED_FIRST_WINS_FACTS: Tuple[str, ...] = (
|
||||
"snapshot_id", "execution_root", "baseline_sha", "target_root",
|
||||
"authority_source", "resource_ref", "selected_subagent_id",
|
||||
"config_fingerprint", "work_order_fingerprint", "work_order_coverage",
|
||||
"authority_fingerprint", "work_order_source_request", "category", "source",
|
||||
*REVIEW_ATTRIBUTION_KEYS)
|
||||
|
|
@ -49,11 +49,22 @@ def _owned_run(ctx: ToolContext, tool: str, run_id: str) -> Tuple[Optional[str],
|
|||
return _fail(tool, "run_ownership_unknown",
|
||||
"No durable record of that run id exists on this drive, so ownership "
|
||||
"cannot be established. Unknown ownership is refused, not waved through.",
|
||||
run_id=run_id), None
|
||||
run_id=run_id,
|
||||
hint="The run may belong to a different drive or the id may be "
|
||||
"mistyped; get_task_result(<task_id>) is the ownership-free "
|
||||
"way to read another task's delegated-run outcome."), None
|
||||
if status == custody.FOREIGN:
|
||||
# The refusal stays a refusal; the additive facts give the caller the
|
||||
# two things it needs to stop being stuck — the run is over, and whom
|
||||
# to ask (get_task_result(owner_task_id) is the legitimate cross-task
|
||||
# read that already carries the delegated_runs_* counters).
|
||||
return _fail(tool, "run_not_owned",
|
||||
"That run belongs to another task. A delegated run may only be "
|
||||
"waited on or cancelled by the task that started it.", run_id=run_id), None
|
||||
"waited on or cancelled by the task that started it.",
|
||||
run_id=run_id,
|
||||
owner_task_id=str(getattr(entry, "task_id", "") or ""),
|
||||
run_settled=bool(getattr(entry, "settled", False)),
|
||||
run_terminal_state=str(getattr(entry, "terminal_state", "") or "")), None
|
||||
return None, entry
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -207,7 +207,8 @@ def _source_range_receipt_valid(
|
|||
def record_started_custody(
|
||||
drive: Any, run_id: str, ctx: Any, route: Any, authority: Any, *,
|
||||
key: str, access: str, root: str, seconds: int, invocation_id: str,
|
||||
project_id: str, project_owned: bool, selected_subagent_id: str,
|
||||
project_id: str, project_owned: bool, project_persistent: bool,
|
||||
selected_subagent_id: str,
|
||||
config_fingerprint: str, work_order_fingerprint: str, work_order_coverage: str,
|
||||
work_order_source_request: Dict[str, Any], authority_fingerprint: str,
|
||||
snapshot_id: str, target_root: str, baseline_sha: str, authority_source: str,
|
||||
|
|
@ -227,6 +228,7 @@ def record_started_custody(
|
|||
profile_id=route.profile_id,
|
||||
project_id=project_id,
|
||||
project_owned=project_owned,
|
||||
project_persistent=project_persistent,
|
||||
root_task_id=str(metadata.get("root_task_id") or ""),
|
||||
parent_task_id=str(metadata.get("parent_task_id") or ""),
|
||||
ledger_root=str(drive),
|
||||
|
|
|
|||
|
|
@ -80,11 +80,23 @@ def claimed_start_request(
|
|||
"Finish it before starting another assignment against the skill."
|
||||
),
|
||||
}
|
||||
live_ids = [
|
||||
str(v) for key in (
|
||||
"open_run_ids", "pending_invocation_ids",
|
||||
"undisposed_patch_run_ids",
|
||||
) for v in (blockers.get(key) or [])
|
||||
]
|
||||
shown = ", ".join(live_ids[:4]) + (
|
||||
f" (+{len(live_ids) - 4} more)" if len(live_ids) > 4 else "")
|
||||
return False, {
|
||||
"reason": "replacement_requires_settlement", **blockers,
|
||||
"detail": (
|
||||
"The actor gained an unsettled start/run or undisposed patch "
|
||||
"before this fresh request could be claimed."
|
||||
"The actor already has an unsettled start/run or an "
|
||||
"undisposed patch"
|
||||
+ (f" ({shown})" if live_ids else "")
|
||||
+ ". Wait for or cancel an open run; replay a pending "
|
||||
"invocation with retry_of=<invocation id>; dispose a "
|
||||
"captured patch explicitly (#364)."
|
||||
),
|
||||
}
|
||||
bootstrap = (
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from typing import Any, Callable, Dict, Mapping, Optional
|
||||
from typing import Any, Callable, Dict, List, Mapping, Optional
|
||||
|
||||
from ouroboros import delegate_custody as custody
|
||||
|
||||
|
|
@ -40,14 +40,42 @@ def terminal_reconcile_task(
|
|||
return result
|
||||
|
||||
|
||||
def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> None:
|
||||
"""Read-only custody audit: fills unreconciled/audit fields and emits evidence."""
|
||||
def custody_audit_snapshot(drive_root: Any) -> Dict[str, Any]:
|
||||
"""One shared read of the custody rows for a BATCH of audits.
|
||||
|
||||
One ``replay()`` pass plus one pending-invocations pass, reused by every
|
||||
audit in the batch — the boot backfill must not rescan the unbounded event
|
||||
log four times per stored row.
|
||||
"""
|
||||
return {
|
||||
"state": custody.replay(drive_root),
|
||||
"pending": custody.pending_invocations(drive_root),
|
||||
}
|
||||
|
||||
|
||||
def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any], *,
|
||||
snapshot: Optional[Mapping[str, Any]] = None,
|
||||
emit_evidence: bool = True) -> None:
|
||||
"""Read-only custody audit: fills unreconciled/audit fields and emits evidence.
|
||||
|
||||
``snapshot`` is a ``custody_audit_snapshot`` shared across a batch of
|
||||
audits; ``emit_evidence=False`` defers the evidence rows so a caller can
|
||||
compare the audit against the stored disclosure first (a no-op refresh must
|
||||
not append custody events every boot).
|
||||
"""
|
||||
state = snapshot.get("state") if snapshot is not None else None
|
||||
pending = snapshot.get("pending") if snapshot is not None else None
|
||||
# The keyword rides only when a snapshot is really shared, so the
|
||||
# no-snapshot call shape stays byte-identical for every existing caller
|
||||
# and test seam over these projections.
|
||||
state_kw: Dict[str, Any] = {} if state is None else {"state": state}
|
||||
audit_failure = ""
|
||||
if custody.custody_log_unreadable(drive_root):
|
||||
audit_failure = "custody_log_unreadable"
|
||||
try:
|
||||
open_ids = (
|
||||
[row.run_id for row in custody.open_runs(drive_root) if row.task_id == mine]
|
||||
[row.run_id for row in custody.open_runs(drive_root, **state_kw)
|
||||
if row.task_id == mine]
|
||||
if not audit_failure else []
|
||||
)
|
||||
except Exception:
|
||||
|
|
@ -56,7 +84,8 @@ def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> N
|
|||
invocation_ids = (
|
||||
[
|
||||
str(row.get("invocation_id") or "")
|
||||
for row in custody.pending_invocations(drive_root)
|
||||
for row in (pending if pending is not None
|
||||
else custody.pending_invocations(drive_root))
|
||||
if str(row.get("task_id") or "") == mine
|
||||
and str(row.get("invocation_id") or "")
|
||||
]
|
||||
|
|
@ -66,7 +95,8 @@ def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> N
|
|||
invocation_ids, audit_failure = [], "pending_invocation_audit_failed"
|
||||
try:
|
||||
patch_ids = (
|
||||
[row.run_id for row in custody.undisposed_patches(drive_root) if row.task_id == mine]
|
||||
[row.run_id for row in custody.undisposed_patches(drive_root, **state_kw)
|
||||
if row.task_id == mine]
|
||||
if not audit_failure else []
|
||||
)
|
||||
except Exception:
|
||||
|
|
@ -76,7 +106,7 @@ def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> N
|
|||
if not audit_failure:
|
||||
deferred_retirements = [
|
||||
row.run_id
|
||||
for row in custody.owned_project_registrations(drive_root)
|
||||
for row in custody.owned_project_registrations(drive_root, **state_kw)
|
||||
if row.task_id == mine and row.settled
|
||||
]
|
||||
except Exception:
|
||||
|
|
@ -103,6 +133,13 @@ def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> N
|
|||
"audit_status": "failed",
|
||||
"unreconciled": [f"delegated_run_state_unknown:{audit_failure}"],
|
||||
})
|
||||
if emit_evidence:
|
||||
_emit_audit_evidence(drive_root, result)
|
||||
|
||||
|
||||
def _emit_audit_evidence(drive_root: Any, result: Mapping[str, Any]) -> None:
|
||||
"""The audit's durable evidence rows — separated so a comparison can run first."""
|
||||
mine = str(result.get("task_id") or "")
|
||||
if result["unreconciled"]:
|
||||
custody.emit(drive_root, "delegated_runs_unreconciled", {
|
||||
"task_id": mine, "trigger": result["trigger"],
|
||||
|
|
@ -120,20 +157,125 @@ def _audit_task_custody(drive_root: Any, mine: str, result: Dict[str, Any]) -> N
|
|||
})
|
||||
|
||||
|
||||
def refresh_terminal_reconciliation(drive_root: Any, task_id: str) -> bool:
|
||||
_EVIDENCE_COUNTER_KEYS = (
|
||||
"delegated_runs_started", "delegated_runs_settled",
|
||||
"delegated_runs_succeeded", "delegated_runs_failed",
|
||||
"delegated_runs_source_unresolved",
|
||||
)
|
||||
|
||||
|
||||
def _stored_evidence_stale(existing: Mapping[str, Any], live: Mapping[str, Any]) -> bool:
|
||||
"""True when the stored CURRENT-TRUTH substrate surfaces disagree with custody.
|
||||
|
||||
Only tasks that ever wrote the harness-dispatch mirror participate: a task
|
||||
with neither top-level counters nor an envelope evidence block was not
|
||||
delegated, and minting one here would fabricate a dispatch record.
|
||||
|
||||
Owner Q2=B x this sprint's 1=A, reconciled by SURFACE: the top-level
|
||||
``delegated_runs_*`` counters are a HISTORICAL SNAPSHOT at the original
|
||||
terminal write and are deliberately NOT compared here — they may honestly
|
||||
read ``settled: 0`` beside a later settlement forever. The current-truth
|
||||
surfaces are what staleness means: the ``subagent_envelope`` evidence
|
||||
mirror, ``actual_substrate``, and ``subscription_cost_usd`` — the fields
|
||||
whose lie (``harness_attempted``/free after a paid successful run) the
|
||||
audit reproduced.
|
||||
"""
|
||||
envelope = existing.get("subagent_envelope")
|
||||
stored_ev = (envelope or {}).get("execution_evidence") if isinstance(envelope, Mapping) else None
|
||||
has_top = any(key in existing for key in _EVIDENCE_COUNTER_KEYS)
|
||||
if not has_top and not isinstance(stored_ev, Mapping):
|
||||
return False
|
||||
if live.get("evidence_read_failed"):
|
||||
# Unreadable custody proves nothing; never rewrite over it.
|
||||
return False
|
||||
|
||||
def _as_int(value: Any) -> int:
|
||||
# A garbage stored counter (str/None) must not raise out into the
|
||||
# caller's blanket except and disable healing forever — treat it as a
|
||||
# mismatch so the row is REWRITTEN to the clean live value.
|
||||
try:
|
||||
return int(value or 0)
|
||||
except (TypeError, ValueError):
|
||||
return -1
|
||||
|
||||
if isinstance(stored_ev, Mapping):
|
||||
for key in _EVIDENCE_COUNTER_KEYS:
|
||||
if _as_int(stored_ev.get(key)) != _as_int(live.get(key)):
|
||||
return True
|
||||
if stored_ev.get("subscription_cost_usd") != live.get("subscription_cost_usd"):
|
||||
return True
|
||||
try:
|
||||
from ouroboros.subagents import actual_substrate
|
||||
|
||||
live_substrate = actual_substrate(live)
|
||||
except Exception:
|
||||
live_substrate = ""
|
||||
if live_substrate and str(existing.get("actual_substrate") or "") not in ("", live_substrate):
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def _rewrite_execution_evidence(drive_root: Any, task_id: str, existing: Mapping[str, Any], live: Mapping[str, Any]) -> None:
|
||||
"""Rewrite the stored envelope evidence + current-truth substrate surfaces
|
||||
from live custody, through the same producers the terminal write used.
|
||||
|
||||
Q2=B split: the top-level ``delegated_runs_*`` counters (and the derived
|
||||
``native_contribution``) stay the historical snapshot — they are filtered
|
||||
OUT of the producer's mirror before the write; only ``actual_substrate``
|
||||
and the envelope evidence (which is where every reader, the executor chip
|
||||
included, takes ``subscription_cost_usd`` from) are healed.
|
||||
"""
|
||||
from ouroboros.subagents import actual_substrate, substrate_result_fields
|
||||
from ouroboros.task_results import STATUS_RUNNING, write_task_result
|
||||
|
||||
envelope = dict(existing.get("subagent_envelope") or {})
|
||||
envelope["execution_evidence"] = dict(live)
|
||||
substrate = actual_substrate(live)
|
||||
if substrate:
|
||||
envelope["actual_substrate"] = substrate
|
||||
mirror = {
|
||||
key: value for key, value in substrate_result_fields(envelope).items()
|
||||
if key not in _EVIDENCE_COUNTER_KEYS and key != "native_contribution"
|
||||
}
|
||||
write_task_result(
|
||||
drive_root, str(task_id or ""),
|
||||
str(existing.get("status") or STATUS_RUNNING),
|
||||
subagent_envelope=envelope,
|
||||
**mirror,
|
||||
)
|
||||
|
||||
|
||||
def refresh_terminal_reconciliation(
|
||||
drive_root: Any, task_id: str, *,
|
||||
trigger: str = "sweep_refresh",
|
||||
snapshot: Optional[Mapping[str, Any]] = None,
|
||||
) -> bool:
|
||||
"""Audit-only refresh of a TERMINAL task's stored custody disclosure.
|
||||
|
||||
The periodic sweep can settle a run AFTER its owning task already wrote its
|
||||
terminal result with a non-empty ``delegated_runs_unreconciled`` — the
|
||||
custody ledger then knows the truth while the stored projection keeps
|
||||
lying (nanny-leaf S1). This re-runs ONLY the read-side audit (never
|
||||
``reconcile_task_runs`` — a refresh must not cancel anything) and rewrites
|
||||
the disclosure through the same recorder, whose guard permits clearing a
|
||||
stale non-empty list. The primary ``reason_code`` is deliberately left
|
||||
untouched (owner Q5=A), and already-rendered chat frames are out of scope —
|
||||
the fixed surfaces are the stored result, task details, and the API view
|
||||
(including retry-lineage projections, which read this row live).
|
||||
"""
|
||||
terminal result — the custody ledger then knows the truth while the stored
|
||||
projection keeps lying (nanny-leaf S1). TWO independent stale classes are
|
||||
healed: a stale ``delegated_runs_unreconciled`` disclosure (audited and
|
||||
re-recorded below), and stored substrate counters/cost that disagree with
|
||||
live custody (``_stored_evidence_stale`` — the PR #402 test pinned only the
|
||||
first class, so ``actual_substrate='harness_attempted'`` and
|
||||
``subscription_cost_usd=None`` survived a successful refresh). This re-runs
|
||||
ONLY the read-side audit (never ``reconcile_task_runs`` — a refresh must
|
||||
not cancel anything) and rewrites through the same recorders. The primary
|
||||
``reason_code`` is deliberately left untouched (owner Q5=A), and
|
||||
already-rendered chat frames are out of scope — the fixed surfaces are the
|
||||
stored result, task details, and the API view (including retry-lineage
|
||||
projections, which read this row live).
|
||||
|
||||
``trigger`` names the refreshing surface on the envelope and its evidence
|
||||
rows (``sweep_refresh`` | ``boot_backfill`` | ``kill_path_clear``);
|
||||
``snapshot`` shares one ``custody_audit_snapshot`` across a batch. An audit
|
||||
that MATCHES the stored disclosure (and finds no stale evidence mirror)
|
||||
performs no write and no emit — a permanently-unreconcilable row must not
|
||||
grow events.jsonl on every boot. Returns True only when a row was really
|
||||
refreshed: the audit evidence is emitted AFTER — and only after — the
|
||||
recorder confirms a changed persisted row, so a lock timeout or a refused
|
||||
write leaves no phantom "refreshed" event behind (R3). """
|
||||
mine = str(task_id or "")
|
||||
if not mine:
|
||||
return False
|
||||
|
|
@ -141,49 +283,278 @@ def refresh_terminal_reconciliation(drive_root: Any, task_id: str) -> bool:
|
|||
from ouroboros.task_results import _TRULY_TERMINAL_STATUSES, load_task_result
|
||||
|
||||
existing = load_task_result(drive_root, mine) or {}
|
||||
if not existing.get("delegated_runs_unreconciled"):
|
||||
return False
|
||||
if str(existing.get("status") or "") not in _TRULY_TERMINAL_STATUSES:
|
||||
return False
|
||||
live = custody.task_execution_evidence(drive_root, mine)
|
||||
evidence_stale = _stored_evidence_stale(existing, live)
|
||||
if not existing.get("delegated_runs_unreconciled") and not evidence_stale:
|
||||
return False
|
||||
except Exception:
|
||||
log.debug("Sweep refresh skipped: task result unreadable for %s", mine, exc_info=True)
|
||||
return False
|
||||
result: Dict[str, Any] = {
|
||||
"task_id": mine, "trigger": "sweep_refresh",
|
||||
"task_id": mine, "trigger": str(trigger or "sweep_refresh"),
|
||||
"outcomes": [], "unreconciled": [], "audit_status": "ok",
|
||||
}
|
||||
_audit_task_custody(drive_root, mine, result)
|
||||
record_terminal_reconciliation(drive_root, mine, result)
|
||||
return True
|
||||
_audit_task_custody(drive_root, mine, result, snapshot=snapshot, emit_evidence=False)
|
||||
if _stored_disclosure_matches(existing, result) and not evidence_stale:
|
||||
return False
|
||||
refreshed = False
|
||||
# Disclosure class: the recorder itself re-checks the no-churn gate and the
|
||||
# monotonic guard; evidence is emitted only after it confirms a landed row.
|
||||
if record_terminal_reconciliation(drive_root, mine, result):
|
||||
_emit_audit_evidence(drive_root, result)
|
||||
refreshed = True
|
||||
# Evidence-mirror class: substrate counters/cost rewritten from live
|
||||
# custody through the same producers the terminal write used.
|
||||
if evidence_stale:
|
||||
try:
|
||||
_rewrite_execution_evidence(drive_root, mine, existing, live)
|
||||
refreshed = True
|
||||
except Exception:
|
||||
log.warning("Sweep evidence rewrite failed for %s", mine, exc_info=True)
|
||||
return refreshed
|
||||
|
||||
|
||||
def backfill_terminal_reconciliations(drive_root: Any) -> List[str]:
|
||||
"""Boot backfill: refresh every stored TERMINAL row that still discloses
|
||||
unreconciled delegated runs.
|
||||
|
||||
The sweep-side refresh covers only task ids named in the CURRENT pass's
|
||||
reconcile outcomes, so a settlement from a previous server generation
|
||||
leaves the stored projection stale forever (generation-crossing residual).
|
||||
Driven by the reverse join — the stored results with a non-empty
|
||||
``delegated_runs_unreconciled`` are a self-clearing set — never by a replay
|
||||
scan of the unbounded event log; one shared snapshot serves every audit,
|
||||
and each row is fail-soft. Returns the task ids actually refreshed.
|
||||
"""
|
||||
try:
|
||||
from ouroboros.task_results import _TRULY_TERMINAL_STATUSES, list_task_results
|
||||
|
||||
stale = [
|
||||
str(row.get("task_id") or "")
|
||||
for row in list_task_results(drive_root)
|
||||
if row.get("delegated_runs_unreconciled")
|
||||
and str(row.get("status") or "") in _TRULY_TERMINAL_STATUSES
|
||||
]
|
||||
except Exception:
|
||||
log.debug("Boot custody-disclosure backfill scan failed", exc_info=True)
|
||||
return []
|
||||
if not stale:
|
||||
return []
|
||||
snapshot = custody_audit_snapshot(drive_root)
|
||||
refreshed: List[str] = []
|
||||
for task_id in stale:
|
||||
try:
|
||||
if refresh_terminal_reconciliation(
|
||||
drive_root, task_id, trigger="boot_backfill", snapshot=snapshot):
|
||||
refreshed.append(task_id)
|
||||
except Exception:
|
||||
log.debug("Boot custody-disclosure backfill failed for %s", task_id, exc_info=True)
|
||||
return refreshed
|
||||
|
||||
|
||||
# The envelope fields whose values change WHAT a reader can conclude about the
|
||||
# task's delegated custody — compared by the no-churn gate below. Provenance
|
||||
# (trigger/outcomes/audit_status) is deliberately excluded: it differs between
|
||||
# otherwise identical audits and would defeat the gate.
|
||||
_ENVELOPE_DISCLOSURE_FIELDS = (
|
||||
"open_run_ids",
|
||||
"pending_invocation_ids",
|
||||
"undisposed_patch_run_ids",
|
||||
"deferred_project_retirements",
|
||||
)
|
||||
|
||||
|
||||
def _stored_disclosure_matches(
|
||||
existing: Mapping[str, Any], result: Mapping[str, Any],
|
||||
) -> bool:
|
||||
"""Whether the stored row already carries exactly this audit's disclosure.
|
||||
|
||||
Compared on the behavior-bearing surfaces (R2): the flat unreconciled list
|
||||
AND every ``_ENVELOPE_DISCLOSURE_FIELDS`` entry of the stored envelope,
|
||||
plus envelope PRESENCE — a kill-written row carrying only the flat list
|
||||
must NOT match, so the next boot adds the envelope once and then becomes a
|
||||
byte-level no-op. The one exception is a row with nothing to disclose at
|
||||
all (no list, no envelope fields): rewriting such a row — or minting one
|
||||
for a task that never delegated — just to attach an empty envelope would
|
||||
be churn, so an absent envelope over an all-empty audit still matches.
|
||||
"""
|
||||
stored_envelope = existing.get("delegate_terminal_reconciliation")
|
||||
has_envelope = isinstance(stored_envelope, dict) and bool(stored_envelope)
|
||||
stored_envelope = stored_envelope if isinstance(stored_envelope, dict) else {}
|
||||
if list(existing.get("delegated_runs_unreconciled") or []) != list(
|
||||
result.get("unreconciled") or []):
|
||||
return False
|
||||
if any(
|
||||
list(stored_envelope.get(field) or []) != list(result.get(field) or [])
|
||||
for field in _ENVELOPE_DISCLOSURE_FIELDS
|
||||
):
|
||||
return False
|
||||
if has_envelope:
|
||||
return True
|
||||
return not (
|
||||
list(result.get("unreconciled") or [])
|
||||
or any(list(result.get(field) or []) for field in _ENVELOPE_DISCLOSURE_FIELDS)
|
||||
)
|
||||
|
||||
|
||||
_REFRESH_CURSOR_REL = "state/delegate_terminal_refresh_cursor.json"
|
||||
_REFRESH_SCAN_CAP_BYTES = 5 * 1024 * 1024 # bounded work per sweep tick
|
||||
_REFRESH_DEFERRED_CAP = 500 # terminal-boundary tasks awaiting their result
|
||||
|
||||
|
||||
def refresh_recently_settled_terminals(drive_root: Any) -> int:
|
||||
"""Refresh terminal results of tasks whose runs settled since the cursor.
|
||||
|
||||
The orphan sweep only revisits tasks named in THIS generation's reconcile
|
||||
outcomes; a run settled at the terminal boundary (or by an earlier
|
||||
generation) never reappears there, so its task's stored evidence stays
|
||||
stale forever. A durable byte-offset cursor over the append-only custody
|
||||
event log keeps each tick bounded to newly appended SETTLED rows (house
|
||||
projection-beside-the-log pattern) — never a full replay per sweep. A
|
||||
shrunken/rotated log resets the cursor; the one-time historical pass is
|
||||
paced by the per-tick byte cap. Returns the number of refreshed tasks.
|
||||
"""
|
||||
import pathlib as _pathlib
|
||||
|
||||
from ouroboros.utils import atomic_write_json, read_json_dict
|
||||
|
||||
log_path = custody.event_log_path(drive_root)
|
||||
cursor_path = _pathlib.Path(drive_root) / _REFRESH_CURSOR_REL
|
||||
try:
|
||||
size = log_path.stat().st_size if log_path.exists() else 0
|
||||
except OSError:
|
||||
return 0
|
||||
stored = read_json_dict(cursor_path) or {}
|
||||
offset = int(stored.get("offset") or 0)
|
||||
if offset > size:
|
||||
offset = 0 # rotated/truncated log: re-ground once
|
||||
deferred: Dict[str, str] = {
|
||||
str(k): str(v) for k, v in (stored.get("deferred") or {}).items() if k
|
||||
}
|
||||
if offset >= size and not deferred:
|
||||
return 0
|
||||
# A settled run whose OWNING TASK has not yet written its terminal result
|
||||
# cannot be healed now (refresh no-ops on a non-terminal task) — the run
|
||||
# settled at the terminal boundary, the exact class this pass exists to
|
||||
# catch. Its task_id goes into the durable ``deferred`` map (retried every
|
||||
# tick) while the byte offset ALWAYS advances: pinning the offset on the
|
||||
# earliest deferred row would re-read the same window forever behind one
|
||||
# long-lived parent and starve every later settlement past the per-tick
|
||||
# byte cap.
|
||||
batch_ids: set = set()
|
||||
end_offset = offset
|
||||
try:
|
||||
with log_path.open("rb") as fh:
|
||||
fh.seek(offset)
|
||||
read_bytes = 0
|
||||
for raw in fh:
|
||||
if not raw.endswith(b"\n") or read_bytes > _REFRESH_SCAN_CAP_BYTES:
|
||||
break # incomplete tail line stays for the next tick
|
||||
read_bytes += len(raw)
|
||||
end_offset += len(raw)
|
||||
row = _json_row(raw)
|
||||
if str(row.get("type") or "") in (custody.SETTLED, custody.CLOSED_ABSENT):
|
||||
tid = str(row.get("task_id") or "")
|
||||
if tid:
|
||||
batch_ids.add(tid)
|
||||
except OSError:
|
||||
return 0
|
||||
from ouroboros.utils import utc_now_iso
|
||||
|
||||
now_iso = utc_now_iso()
|
||||
refreshed = 0
|
||||
next_deferred: Dict[str, str] = {}
|
||||
for tid in sorted(batch_ids | set(deferred)):
|
||||
since = deferred.get(tid) or now_iso
|
||||
try:
|
||||
if _task_is_terminal(drive_root, tid):
|
||||
if refresh_terminal_reconciliation(drive_root, tid):
|
||||
refreshed += 1
|
||||
else:
|
||||
next_deferred[tid] = since
|
||||
except Exception:
|
||||
log.debug("Cursor refresh failed for %s", tid, exc_info=True)
|
||||
next_deferred[tid] = since
|
||||
if len(next_deferred) > _REFRESH_DEFERRED_CAP:
|
||||
# Oldest-first eviction, disclosed: a task deferred this long with no
|
||||
# terminal result is overwhelmingly abandoned; keeping the map bounded
|
||||
# protects the cursor write from unbounded growth.
|
||||
keep = sorted(next_deferred.items(), key=lambda kv: kv[1])[-_REFRESH_DEFERRED_CAP:]
|
||||
dropped = len(next_deferred) - len(keep)
|
||||
next_deferred = dict(keep)
|
||||
log.info("Settled-refresh deferred map over cap; dropped %d oldest entries", dropped)
|
||||
try:
|
||||
atomic_write_json(cursor_path, {"offset": end_offset, "deferred": next_deferred})
|
||||
except Exception:
|
||||
log.debug("Refresh cursor write failed", exc_info=True)
|
||||
return refreshed
|
||||
|
||||
|
||||
def _json_row(raw: bytes) -> Dict[str, Any]:
|
||||
import json as _json
|
||||
|
||||
try:
|
||||
row = _json.loads(raw)
|
||||
return row if isinstance(row, dict) else {}
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
|
||||
def _task_is_terminal(drive_root: Any, task_id: str) -> bool:
|
||||
"""Whether the task already wrote a truly-terminal result (heal-eligible)."""
|
||||
try:
|
||||
from ouroboros.task_results import _TRULY_TERMINAL_STATUSES, load_task_result
|
||||
|
||||
existing = load_task_result(drive_root, str(task_id or "")) or {}
|
||||
return str(existing.get("status") or "") in _TRULY_TERMINAL_STATUSES
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
def record_terminal_reconciliation(
|
||||
drive_root: Any, task_id: str, result: Mapping[str, Any],
|
||||
) -> None:
|
||||
"""Attach the audit to the task record without choosing lifecycle policy."""
|
||||
) -> bool:
|
||||
"""Attach the audit to the task record without choosing lifecycle policy.
|
||||
|
||||
Returns True only when the row was actually persisted AND now carries this
|
||||
audit's disclosure (R3) — a lock timeout, a raising store, or a write the
|
||||
monotonic guard refused all report False, so callers (the guarded refresh,
|
||||
the boot backfill) cannot claim a heal that never landed.
|
||||
"""
|
||||
|
||||
try:
|
||||
from ouroboros.task_results import STATUS_RUNNING, load_task_result, write_task_result
|
||||
|
||||
existing = load_task_result(drive_root, str(task_id or "")) or {}
|
||||
unreconciled = list(result.get("unreconciled") or [])
|
||||
deferred = list(result.get("deferred_project_retirements") or [])
|
||||
if (not unreconciled and not deferred
|
||||
and not existing.get("delegated_runs_unreconciled")):
|
||||
return
|
||||
write_task_result(
|
||||
# Write only when the disclosure MOVES. Subsumes the old empty-over-empty
|
||||
# guard: a task that never delegated still gets no row minted here (the
|
||||
# STATUS_RUNNING fallback below exists for a legitimately mid-flight
|
||||
# loop-exit row, never for inventing one), and an already-current
|
||||
# disclosure is not rewritten on every kill or boot.
|
||||
if _stored_disclosure_matches(existing, result):
|
||||
return False
|
||||
stored = write_task_result(
|
||||
drive_root,
|
||||
str(task_id or ""),
|
||||
str(existing.get("status") or STATUS_RUNNING),
|
||||
delegate_terminal_reconciliation=dict(result),
|
||||
delegated_runs_unreconciled=unreconciled,
|
||||
delegated_runs_unreconciled=list(result.get("unreconciled") or []),
|
||||
)
|
||||
return _stored_disclosure_matches(
|
||||
stored if isinstance(stored, dict) else {}, result,
|
||||
)
|
||||
except Exception:
|
||||
log.warning("Failed to persist terminal custody audit for %s", task_id, exc_info=True)
|
||||
return False
|
||||
|
||||
|
||||
__all__ = [
|
||||
"backfill_terminal_reconciliations",
|
||||
"custody_audit_snapshot",
|
||||
"record_terminal_reconciliation",
|
||||
"refresh_recently_settled_terminals",
|
||||
"refresh_terminal_reconciliation",
|
||||
"terminal_reconcile_task",
|
||||
]
|
||||
|
|
|
|||
|
|
@ -17,8 +17,10 @@ CHAT_TYPING = "chat.typing"
|
|||
CHAT_PHOTO = "chat.photo"
|
||||
CHAT_VIDEO = "chat.video"
|
||||
CHAT_DOCUMENT = "chat.document"
|
||||
CHAT_LINKS = "chat.links"
|
||||
CHAT_QUIZ = "chat.quiz"
|
||||
SKILL_LIFECYCLE = "skill.lifecycle"
|
||||
VALID_TOPICS = frozenset({CHAT_OUTBOUND, CHAT_TYPING, CHAT_PHOTO, CHAT_VIDEO, CHAT_DOCUMENT, SKILL_LIFECYCLE})
|
||||
VALID_TOPICS = frozenset({CHAT_OUTBOUND, CHAT_TYPING, CHAT_PHOTO, CHAT_VIDEO, CHAT_DOCUMENT, CHAT_LINKS, CHAT_QUIZ, SKILL_LIFECYCLE})
|
||||
|
||||
|
||||
@dataclass
|
||||
|
|
@ -130,8 +132,10 @@ def publish_event(topic: str, data: Dict[str, Any]) -> None:
|
|||
|
||||
__all__ = [
|
||||
"CHAT_DOCUMENT",
|
||||
"CHAT_LINKS",
|
||||
"CHAT_OUTBOUND",
|
||||
"CHAT_PHOTO",
|
||||
"CHAT_QUIZ",
|
||||
"CHAT_TYPING",
|
||||
"CHAT_VIDEO",
|
||||
"EventBus",
|
||||
|
|
|
|||
|
|
@ -204,11 +204,19 @@ def materialize_companion_env(api: "PluginAPIImpl", descriptor: Any, spec: Dict[
|
|||
granted_keys=list(api._granted_upper),
|
||||
)
|
||||
reserved_env = {"HOST_SERVICE_TOKEN", "HOST_SERVICE_URL"}
|
||||
for key, value in (spec.get("env") or {}).items():
|
||||
key_text = str(key)
|
||||
if key_text.upper() in FORBIDDEN_SKILL_SETTINGS or key_text.upper() in reserved_env:
|
||||
continue
|
||||
env[key_text] = str(value)
|
||||
manifest_env = {
|
||||
str(key): str(value)
|
||||
for key, value in (spec.get("env") or {}).items()
|
||||
if str(key).upper() not in FORBIDDEN_SKILL_SETTINGS
|
||||
and str(key).upper() not in reserved_env
|
||||
}
|
||||
# Case-aware merge (delta finding D2-8): a manifest "Path" must REPLACE
|
||||
# the allowlisted "PATH" on Windows, never sit next to it — duplicate
|
||||
# case-variant env keys make CreateProcess-era spawns fail or pick an
|
||||
# undefined winner. Same contract as the executor-local service lane.
|
||||
from ouroboros.workspace_executor import overlay_env
|
||||
|
||||
env = overlay_env(env, manifest_env)
|
||||
env["HOST_SERVICE_URL"] = f"http://{DEFAULT_HOST_SERVICE_HOST}:{host_service_port()}"
|
||||
env["HOST_SERVICE_TOKEN"] = token
|
||||
if api._skill_dir is not None:
|
||||
|
|
@ -218,5 +226,24 @@ def materialize_companion_env(api: "PluginAPIImpl", descriptor: Any, spec: Dict[
|
|||
env["PYTHONPATH"] = os.pathsep.join(
|
||||
[*site_dirs, existing_pythonpath] if existing_pythonpath else site_dirs
|
||||
)
|
||||
expected_runtime = str(spec.get("runtime") or "").strip()
|
||||
manifest_path_override = any(
|
||||
str(key).upper() == "PATH" for key in (spec.get("env") or {})
|
||||
)
|
||||
if expected_runtime in {"node", "npm"} and not manifest_path_override:
|
||||
# T14 emergency-only PATH prepend (see register_companion_process):
|
||||
# descriptor env keys win over the supervisor's `_companion_base_env`
|
||||
# merge, so the prepended PATH reaches the child and survives
|
||||
# supervisor restarts.
|
||||
from ouroboros.platform_layer import skill_node_emergency_path_dir
|
||||
|
||||
node_prepend_dir = skill_node_emergency_path_dir()
|
||||
if node_prepend_dir:
|
||||
existing_path = env.get("PATH") or os.environ.get("PATH", "")
|
||||
env["PATH"] = (
|
||||
os.pathsep.join([node_prepend_dir, existing_path])
|
||||
if existing_path
|
||||
else node_prepend_dir
|
||||
)
|
||||
descriptor.env.clear()
|
||||
descriptor.env.update(env)
|
||||
|
|
|
|||
|
|
@ -556,6 +556,32 @@ class PluginAPIImpl:
|
|||
raise ExtensionRegistrationError("companion command must be declared in manifest")
|
||||
if expected_runtime in {"python", "python3"} and cmd[0] in {"python", "python3"}:
|
||||
cmd = [sys.executable, *cmd[1:]]
|
||||
manifest_path_override = any(
|
||||
str(key).upper() == "PATH" for key in (spec.get("env") or {})
|
||||
)
|
||||
if expected_runtime in {"node", "npm"} and not manifest_path_override:
|
||||
# T14: an explicit PATH in the companion manifest means the author
|
||||
# owns the runtime lookup — the host-side probe cannot see that
|
||||
# child-only PATH, so neither the bundled rewrite nor the emergency
|
||||
# prepend may shadow it.
|
||||
# Node symmetry with the python->sys.executable rewrite above: the
|
||||
# skill-family runtime precedence (bundled-first + health rollback
|
||||
# to a working PATH node) is owned by
|
||||
# platform_layer.select_skill_node_runtime. npm itself is NOT
|
||||
# bundled and is never rewritten; only in the emergency state
|
||||
# (PATH node missing/execution-probed broken, healthy bundled
|
||||
# selected) does the companion child PATH gain the bundled-node
|
||||
# dir so npm's `#!/usr/bin/env node` shebang finds a working
|
||||
# runtime (the PATH prepend itself happens in the post-fence env
|
||||
# materialization, ouroboros/extension_child_catalog.py). On a
|
||||
# healthy PATH the child env stays byte-identical. Disclosed
|
||||
# residual: an npm launcher rewritten to an ABSOLUTE node shebang
|
||||
# ignores PATH and keeps failing honestly.
|
||||
from ouroboros.platform_layer import select_skill_node_runtime
|
||||
|
||||
selected_node, _node_provenance = select_skill_node_runtime()
|
||||
if selected_node and cmd[0] == "node":
|
||||
cmd = [selected_node, *cmd[1:]]
|
||||
if not is_server_process():
|
||||
with _lock:
|
||||
self._require_open_locked()
|
||||
|
|
|
|||
|
|
@ -112,6 +112,11 @@ class ChatOutbound(TypedDict):
|
|||
# while post-task synthesis still runs, so the frame is NOT the task's
|
||||
# terminal conclusion — task_done settles the card/turn.
|
||||
task_phase: NotRequired[str]
|
||||
# Typed terminal fact on a frame that IS the turn's conclusion: stamped on
|
||||
# direct/ephemeral finals (and the direct error branch) so the client's
|
||||
# live gate settles the activity without waiting for a snapshot. One of
|
||||
# completed/failed/cancelled/rejected_duplicate.
|
||||
task_terminal_status: NotRequired[str]
|
||||
ephemeral_decision: NotRequired[bool]
|
||||
task_incident: NotRequired[str]
|
||||
toast_once: NotRequired[str]
|
||||
|
|
@ -234,6 +239,7 @@ class PhotoOutbound(TypedDict):
|
|||
client_message_id: NotRequired[str]
|
||||
transport: NotRequired[TransportMetadata]
|
||||
chat_id: NotRequired[int]
|
||||
task_id: NotRequired[str]
|
||||
# Server-stamped when chat_id is a reserved Project thread: Main never
|
||||
# adopts it, even before the browser has learned the project.
|
||||
project_thread: NotRequired[bool]
|
||||
|
|
@ -255,6 +261,7 @@ class VideoOutbound(TypedDict):
|
|||
client_message_id: NotRequired[str]
|
||||
transport: NotRequired[TransportMetadata]
|
||||
chat_id: NotRequired[int]
|
||||
task_id: NotRequired[str]
|
||||
# Server-stamped when chat_id is a reserved Project thread: Main never
|
||||
# adopts it, even before the browser has learned the project.
|
||||
project_thread: NotRequired[bool]
|
||||
|
|
@ -281,11 +288,83 @@ class DocumentOutbound(TypedDict):
|
|||
client_message_id: NotRequired[str]
|
||||
transport: NotRequired[TransportMetadata]
|
||||
chat_id: NotRequired[int]
|
||||
task_id: NotRequired[str]
|
||||
size_bytes: NotRequired[int]
|
||||
# Server-stamped when chat_id is a reserved Project thread: Main never
|
||||
# adopts it, even before the browser has learned the project.
|
||||
project_thread: NotRequired[bool]
|
||||
|
||||
|
||||
class LinkAction(TypedDict):
|
||||
"""One validated HTTP(S) action in a structured links frame."""
|
||||
|
||||
label: str
|
||||
url: str
|
||||
|
||||
|
||||
class LinksOutbound(TypedDict):
|
||||
"""Outbound group of first-class external link buttons."""
|
||||
|
||||
type: Literal["links"]
|
||||
role: Literal["assistant"]
|
||||
actions: list[LinkAction]
|
||||
ts: str
|
||||
title: NotRequired[str]
|
||||
chat_id: NotRequired[int]
|
||||
task_id: NotRequired[str]
|
||||
project_thread: NotRequired[bool]
|
||||
transport: NotRequired[TransportMetadata]
|
||||
|
||||
|
||||
class QuizOption(TypedDict):
|
||||
"""One selectable option on an owner quiz card."""
|
||||
|
||||
label: str
|
||||
detail: NotRequired[str]
|
||||
|
||||
|
||||
class QuizOutbound(TypedDict):
|
||||
"""Outbound owner quiz card: a typed question with option buttons.
|
||||
|
||||
Fire-and-continue: the asking task keeps working under ``assumption``
|
||||
while the card is open. ``state`` is the card's lifecycle word
|
||||
(``open`` in this display phase; answered/expired arrive with the
|
||||
answer ingress).
|
||||
"""
|
||||
|
||||
type: Literal["quiz"]
|
||||
role: Literal["assistant"]
|
||||
quiz_id: str
|
||||
question: str
|
||||
options: list[QuizOption]
|
||||
stake: str
|
||||
assumption: str
|
||||
state: str
|
||||
ts: str
|
||||
chat_id: NotRequired[int]
|
||||
task_id: NotRequired[str]
|
||||
project_thread: NotRequired[bool]
|
||||
transport: NotRequired[TransportMetadata]
|
||||
|
||||
|
||||
class QuizStateOutbound(TypedDict):
|
||||
"""Outbound WS lifecycle update for an already-rendered quiz card.
|
||||
|
||||
A separate discriminator (not a second ``quiz`` frame): the display path
|
||||
dedupes quiz frames by ``quiz:{quiz_id}:{ts}``, so a state change must
|
||||
never look like a new card. ``answered_index`` rides only with the
|
||||
``answered`` state.
|
||||
"""
|
||||
|
||||
type: Literal["quiz_state"]
|
||||
quiz_id: str
|
||||
task_id: str
|
||||
state: str
|
||||
ts: str
|
||||
answered_index: NotRequired[int]
|
||||
chat_id: NotRequired[int]
|
||||
|
||||
|
||||
class TypingOutbound(TypedDict):
|
||||
"""Outbound WS typing indicator."""
|
||||
|
||||
|
|
@ -362,6 +441,10 @@ class MessageAnnotationOutbound(TypedDict):
|
|||
target_label: NotRequired[str]
|
||||
options: NotRequired[List[Dict[str, Any]]]
|
||||
attachment_manifest: NotRequired[List[AttachmentManifestEntry]]
|
||||
# #198: the exact refusal-attempt identity — the picker card composes its
|
||||
# decision_id (routing:{client_message_id}:{routing_token}) from it; a
|
||||
# presentation frame without it renders text, never a clickable card.
|
||||
routing_token: NotRequired[str]
|
||||
ts: NotRequired[str]
|
||||
|
||||
|
||||
|
|
@ -569,8 +652,10 @@ class ActiveChatActivity(TypedDict):
|
|||
|
||||
The combined snapshot: direct/ephemeral registry turns (same rows as
|
||||
``active_direct_turns``) plus ROOT managed queue tasks projected as
|
||||
``kind="managed_task"`` with ``phase`` ``queued`` | ``working`` |
|
||||
``finalizing`` (final answer stored, post-task synthesis still open).
|
||||
``kind="managed_task"`` with ``phase`` ``queued`` | ``budget_paused``
|
||||
(zero-dispatch member awaiting an explicit resume — never plain
|
||||
"queued") | ``working`` | ``finalizing`` (final answer stored, post-task
|
||||
synthesis still open).
|
||||
Field shape mirrors ``ActiveDirectTurn`` so one client reducer hydrates
|
||||
both; managed rows carry an empty ``client_message_id``.
|
||||
"""
|
||||
|
|
@ -744,7 +829,10 @@ class UiPreferencesResponse(TypedDict):
|
|||
|
||||
class GitLogResponse(TypedDict):
|
||||
commits: list[Dict[str, Any]]
|
||||
tags: list[str]
|
||||
# Tag rows: {tag, date, sha (peeled commit), message} — the mirror said
|
||||
# ``list[str]`` while ``list_versions`` has always emitted dicts; corrected
|
||||
# (behavioural documentation) in the 2026-08-31 updates redesign.
|
||||
tags: list[Dict[str, Any]]
|
||||
branch: str
|
||||
sha: str
|
||||
|
||||
|
|
@ -1190,6 +1278,44 @@ class TaskHurryResponse(TypedDict, total=False):
|
|||
error: str
|
||||
|
||||
|
||||
class DecisionRequest(TypedDict):
|
||||
"""POST /api/decisions body — the ONE answer ingress for owner decision
|
||||
cards (owner decision 1=A). ``decision_id`` is a composed family id:
|
||||
``quiz:{task_id}:{quiz_id}`` (this phase), ``routing:{client_message_id}:
|
||||
{routing_token}`` (#198), ``interaction:{task_id}:{run_id}:
|
||||
{interaction_id}`` (#204). ``request_id`` is the idempotency key; a
|
||||
replayed request returns the recorded confirmation instead of acting
|
||||
twice. ``comment`` is the owner's optional verbatim remark."""
|
||||
|
||||
request_id: str
|
||||
decision_id: str
|
||||
option_index: int
|
||||
comment: NotRequired[str]
|
||||
|
||||
|
||||
class DecisionResponse(TypedDict, total=False):
|
||||
"""Answer-ingress reply. 2xx carries the card's new lifecycle ``state``
|
||||
(``answered``; ``duplicate`` marks an idempotent replay). A late answer
|
||||
to a settled task is 409 with ``state`` telling the truth
|
||||
(``expired_terminal``/``answered``) so the card settles instead of
|
||||
inviting retries. The routing family (#198) adds: ``dispatched`` (the
|
||||
confirmed durable receipt status), ``task_id`` (the derived id of a
|
||||
promoted task), ``latest_status`` (the superseding row's status on a 409),
|
||||
``reason``/``detail`` (typed refusal/unconfirmed diagnostics)."""
|
||||
|
||||
ok: bool
|
||||
decision_id: str
|
||||
state: str
|
||||
answered_index: int
|
||||
duplicate: bool
|
||||
error: str
|
||||
dispatched: str
|
||||
task_id: str
|
||||
latest_status: str
|
||||
reason: str
|
||||
detail: str
|
||||
|
||||
|
||||
class LogTailResponse(TypedDict, total=False):
|
||||
name: str
|
||||
entries: list[Dict[str, Any]]
|
||||
|
|
@ -1280,117 +1406,7 @@ class OnboardingPresetFailureResponse(TypedDict):
|
|||
|
||||
|
||||
# Human/test-visible contract index; routers own executable Route objects.
|
||||
HTTP_ENDPOINTS: tuple[str, ...] = (
|
||||
"GET /api/health",
|
||||
"GET /api/state",
|
||||
"GET /api/settings",
|
||||
"POST /api/settings",
|
||||
"GET /api/ui/preferences",
|
||||
"POST /api/ui/preferences",
|
||||
"POST /api/owner/runtime-mode",
|
||||
"POST /api/owner/auto-grant",
|
||||
"POST /api/owner/context-mode",
|
||||
"POST /api/owner/safety-mode",
|
||||
"POST /api/owner/capability-ack",
|
||||
"POST /api/owner/skills/{skill}/attest-review",
|
||||
"POST /api/owner/skills/{skill}/presence-runtime",
|
||||
"POST /api/skills/{skill}/publish-preflight",
|
||||
"GET /api/model-catalog",
|
||||
"POST /api/tasks",
|
||||
"GET /api/tasks",
|
||||
"GET /api/tasks/{task_id}",
|
||||
"GET /api/tasks/{task_id}/artifacts/{name}",
|
||||
"GET /api/tasks/{task_id}/events",
|
||||
"POST /api/tasks/{task_id}/cancel",
|
||||
"POST /api/tasks/{task_id}/hurry",
|
||||
"POST /api/tasks/{task_id}/resume",
|
||||
"GET /api/schedules",
|
||||
"POST /api/schedules",
|
||||
"DELETE /api/schedules/{schedule_id}",
|
||||
"POST /api/command",
|
||||
"POST /api/reset",
|
||||
"GET /api/git/log",
|
||||
"POST /api/git/rollback",
|
||||
"POST /api/git/promote",
|
||||
"GET /api/update/status",
|
||||
"POST /api/update/check",
|
||||
"POST /api/update/preflight",
|
||||
"POST /api/update/apply",
|
||||
"GET /api/cost-breakdown",
|
||||
"GET /api/evolution-data",
|
||||
"GET /api/projects",
|
||||
"POST /api/projects",
|
||||
"POST /api/projects/from-task",
|
||||
"POST /api/projects/{project_id}/update",
|
||||
"POST /api/projects/{project_id}/delete",
|
||||
"GET /api/fs/dirs",
|
||||
"GET /api/chat/history",
|
||||
"GET /api/logs/{name}",
|
||||
"POST /api/chat/upload",
|
||||
"DELETE /api/chat/upload",
|
||||
"POST /api/openai-compatible/models",
|
||||
"POST /api/providers/test",
|
||||
"GET /api/local-model/status",
|
||||
"POST /api/local-model/start",
|
||||
"POST /api/local-model/stop",
|
||||
"POST /api/local-model/test",
|
||||
"POST /api/local-model/install-runtime",
|
||||
"GET /api/mcp/status",
|
||||
"POST /api/mcp/refresh",
|
||||
"POST /api/mcp/test",
|
||||
"GET /api/reviewer-slots",
|
||||
"GET /api/claudexor/status",
|
||||
"POST /api/claudexor/wake",
|
||||
"POST /api/claudexor/login",
|
||||
"GET /api/claudexor/login/{job_id}",
|
||||
"DELETE /api/claudexor/login/{job_id}",
|
||||
"POST /api/claudexor/login/{job_id}/input",
|
||||
"POST /api/claudexor/login/{job_id}/reconcile",
|
||||
"DELETE /api/claudexor/credential-profiles/{harness}/{profile_id}",
|
||||
"PATCH /api/claudexor/credential-profiles/{harness}/{profile_id}",
|
||||
"GET /api/extensions",
|
||||
"GET /api/extensions/{skill}/manifest",
|
||||
"GET /api/extensions/{skill}/module/{entry}",
|
||||
"GET /api/extensions/{skill}/settings_section",
|
||||
"ANY /api/extensions/{skill}/{rest:path}",
|
||||
"GET /api/skills/daemons",
|
||||
"POST /api/skills/{skill}/toggle",
|
||||
"POST /api/skills/{skill}/delete",
|
||||
"GET /api/skills/lifecycle-queue",
|
||||
"POST /api/skills/{skill}/review",
|
||||
"GET /api/skills/{skill}/review-history/{job_id}",
|
||||
"POST /api/skills/{skill}/grants",
|
||||
"POST /api/skills/{skill}/reconcile",
|
||||
"GET /api/marketplace/clawhub/search",
|
||||
"GET /api/marketplace/clawhub/installed",
|
||||
"GET /api/marketplace/clawhub/info/{slug:path}",
|
||||
"GET /api/marketplace/clawhub/preview/{slug:path}",
|
||||
"POST /api/marketplace/clawhub/install",
|
||||
"POST /api/marketplace/clawhub/update/{name}",
|
||||
"POST /api/marketplace/clawhub/uninstall/{name}",
|
||||
"GET /api/marketplace/ouroboroshub/catalog",
|
||||
"GET /api/marketplace/ouroboroshub/installed",
|
||||
"GET /api/marketplace/ouroboroshub/preview/{slug:path}",
|
||||
"POST /api/marketplace/ouroboroshub/install",
|
||||
"POST /api/marketplace/ouroboroshub/update/{name}",
|
||||
"POST /api/marketplace/ouroboroshub/uninstall/{name}",
|
||||
# The wizard PAGE (one onboarding host: desktop setup window, blocking
|
||||
# overlay frame, plain browser). /api/onboarding stays the readiness probe.
|
||||
"GET /onboarding",
|
||||
"GET /api/onboarding",
|
||||
"POST /api/onboarding/subagents/preview",
|
||||
"POST /api/onboarding/complete",
|
||||
"GET /api/files/list",
|
||||
"GET /api/files/read",
|
||||
"GET /api/files/content",
|
||||
"GET /api/files/download",
|
||||
"POST /api/files/upload",
|
||||
"POST /api/files/mkdir",
|
||||
"POST /api/files/write",
|
||||
"POST /api/files/delete",
|
||||
"POST /api/files/transfer",
|
||||
"WS /ws",
|
||||
)
|
||||
from ouroboros.gateway.endpoint_index import HTTP_ENDPOINTS
|
||||
|
||||
WS_MESSAGE_TYPES: tuple[str, ...] = (
|
||||
"chat",
|
||||
|
|
@ -1398,6 +1414,9 @@ WS_MESSAGE_TYPES: tuple[str, ...] = (
|
|||
"photo",
|
||||
"video",
|
||||
"document",
|
||||
"links",
|
||||
"quiz",
|
||||
"quiz_state",
|
||||
"typing",
|
||||
"log",
|
||||
"heartbeat",
|
||||
|
|
@ -1419,6 +1438,13 @@ __all__ = [
|
|||
"PhotoOutbound",
|
||||
"VideoOutbound",
|
||||
"DocumentOutbound",
|
||||
"LinkAction",
|
||||
"LinksOutbound",
|
||||
"QuizOption",
|
||||
"QuizOutbound",
|
||||
"QuizStateOutbound",
|
||||
"DecisionRequest",
|
||||
"DecisionResponse",
|
||||
"TypingOutbound",
|
||||
"LogOutbound",
|
||||
"HeartbeatOutbound",
|
||||
|
|
|
|||
|
|
@ -46,8 +46,26 @@ def _runtime_branch_defaults(request: Request) -> tuple[str, str]:
|
|||
|
||||
def _managed_update_payload(*, fetch: bool, include_tags: bool) -> dict[str, Any]:
|
||||
from supervisor.git_ops import compute_managed_update_status, git_capture
|
||||
from supervisor.update_merge import active_update_tx
|
||||
|
||||
status = compute_managed_update_status(fetch=fetch)
|
||||
# Additive minimal public projection of an active managed-update
|
||||
# transaction, so a re-opened panel can say "resolution in progress"
|
||||
# instead of silently reading as ordinary state (a second apply 409s).
|
||||
try:
|
||||
tx = active_update_tx()
|
||||
except Exception:
|
||||
tx = {}
|
||||
update_tx = (
|
||||
{
|
||||
"active": True,
|
||||
"phase": str(tx.get("phase") or ""),
|
||||
"task_id": str(tx.get("task_id") or ""),
|
||||
"restart_required": bool(tx.get("restart_required")),
|
||||
}
|
||||
if tx
|
||||
else {"active": False}
|
||||
)
|
||||
latest_version = ""
|
||||
latest_sha = status.get("latest_sha") or ""
|
||||
if latest_sha:
|
||||
|
|
@ -63,6 +81,7 @@ def _managed_update_payload(*, fetch: bool, include_tags: bool) -> dict[str, Any
|
|||
"current_version": get_version(),
|
||||
"latest_version": latest_version,
|
||||
"official_tags": official_tags,
|
||||
"update_tx": update_tx,
|
||||
**status,
|
||||
}
|
||||
|
||||
|
|
|
|||
122
ouroboros/gateway/endpoint_index.py
Normal file
122
ouroboros/gateway/endpoint_index.py
Normal file
|
|
@ -0,0 +1,122 @@
|
|||
"""Human/test-visible HTTP endpoint index for the browser gateway ABI.
|
||||
|
||||
Routers own the executable Route objects; this tuple is the reviewable
|
||||
contract index consumed by the parity/reliability suites. It lives beside
|
||||
``gateway/contracts.py`` (which re-exports it) so the mirror file keeps its
|
||||
TypedDict envelopes within the module size gate.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
HTTP_ENDPOINTS: tuple[str, ...] = (
|
||||
"GET /api/health",
|
||||
"GET /api/state",
|
||||
"GET /api/settings",
|
||||
"POST /api/settings",
|
||||
"GET /api/ui/preferences",
|
||||
"POST /api/ui/preferences",
|
||||
"POST /api/owner/runtime-mode",
|
||||
"POST /api/owner/auto-grant",
|
||||
"POST /api/owner/context-mode",
|
||||
"POST /api/owner/safety-mode",
|
||||
"POST /api/owner/capability-ack",
|
||||
"POST /api/owner/skills/{skill}/attest-review",
|
||||
"POST /api/owner/skills/{skill}/presence-runtime",
|
||||
"POST /api/skills/{skill}/publish-preflight",
|
||||
"GET /api/model-catalog",
|
||||
"POST /api/tasks",
|
||||
"GET /api/tasks",
|
||||
"GET /api/tasks/{task_id}",
|
||||
"GET /api/tasks/{task_id}/artifacts/{name}",
|
||||
"GET /api/tasks/{task_id}/events",
|
||||
"POST /api/tasks/{task_id}/cancel",
|
||||
"POST /api/tasks/{task_id}/hurry",
|
||||
"POST /api/tasks/{task_id}/resume",
|
||||
"POST /api/decisions",
|
||||
"GET /api/schedules",
|
||||
"POST /api/schedules",
|
||||
"DELETE /api/schedules/{schedule_id}",
|
||||
"POST /api/command",
|
||||
"POST /api/reset",
|
||||
"GET /api/git/log",
|
||||
"POST /api/git/rollback",
|
||||
"POST /api/git/promote",
|
||||
"GET /api/update/status",
|
||||
"POST /api/update/check",
|
||||
"POST /api/update/preflight",
|
||||
"POST /api/update/apply",
|
||||
"GET /api/cost-breakdown",
|
||||
"GET /api/evolution-data",
|
||||
"GET /api/projects",
|
||||
"POST /api/projects",
|
||||
"POST /api/projects/from-task",
|
||||
"POST /api/projects/{project_id}/update",
|
||||
"POST /api/projects/{project_id}/delete",
|
||||
"GET /api/fs/dirs",
|
||||
"GET /api/chat/history",
|
||||
"GET /api/logs/{name}",
|
||||
"POST /api/chat/upload",
|
||||
"DELETE /api/chat/upload",
|
||||
"POST /api/openai-compatible/models",
|
||||
"POST /api/providers/test",
|
||||
"GET /api/local-model/status",
|
||||
"POST /api/local-model/start",
|
||||
"POST /api/local-model/stop",
|
||||
"POST /api/local-model/test",
|
||||
"POST /api/local-model/install-runtime",
|
||||
"GET /api/mcp/status",
|
||||
"POST /api/mcp/refresh",
|
||||
"POST /api/mcp/test",
|
||||
"GET /api/reviewer-slots",
|
||||
"GET /api/claudexor/status",
|
||||
"POST /api/claudexor/wake",
|
||||
"POST /api/claudexor/login",
|
||||
"GET /api/claudexor/login/{job_id}",
|
||||
"DELETE /api/claudexor/login/{job_id}",
|
||||
"POST /api/claudexor/login/{job_id}/input",
|
||||
"POST /api/claudexor/login/{job_id}/reconcile",
|
||||
"DELETE /api/claudexor/credential-profiles/{harness}/{profile_id}",
|
||||
"PATCH /api/claudexor/credential-profiles/{harness}/{profile_id}",
|
||||
"GET /api/extensions",
|
||||
"GET /api/extensions/{skill}/manifest",
|
||||
"GET /api/extensions/{skill}/module/{entry}",
|
||||
"GET /api/extensions/{skill}/settings_section",
|
||||
"ANY /api/extensions/{skill}/{rest:path}",
|
||||
"GET /api/skills/daemons",
|
||||
"POST /api/skills/{skill}/toggle",
|
||||
"POST /api/skills/{skill}/delete",
|
||||
"GET /api/skills/lifecycle-queue",
|
||||
"POST /api/skills/{skill}/review",
|
||||
"GET /api/skills/{skill}/review-history/{job_id}",
|
||||
"POST /api/skills/{skill}/grants",
|
||||
"POST /api/skills/{skill}/reconcile",
|
||||
"GET /api/marketplace/clawhub/search",
|
||||
"GET /api/marketplace/clawhub/installed",
|
||||
"GET /api/marketplace/clawhub/info/{slug:path}",
|
||||
"GET /api/marketplace/clawhub/preview/{slug:path}",
|
||||
"POST /api/marketplace/clawhub/install",
|
||||
"POST /api/marketplace/clawhub/update/{name}",
|
||||
"POST /api/marketplace/clawhub/uninstall/{name}",
|
||||
"GET /api/marketplace/ouroboroshub/catalog",
|
||||
"GET /api/marketplace/ouroboroshub/installed",
|
||||
"GET /api/marketplace/ouroboroshub/preview/{slug:path}",
|
||||
"POST /api/marketplace/ouroboroshub/install",
|
||||
"POST /api/marketplace/ouroboroshub/update/{name}",
|
||||
"POST /api/marketplace/ouroboroshub/uninstall/{name}",
|
||||
# The wizard PAGE (one onboarding host: desktop setup window, blocking
|
||||
# overlay frame, plain browser). /api/onboarding stays the readiness probe.
|
||||
"GET /onboarding",
|
||||
"GET /api/onboarding",
|
||||
"POST /api/onboarding/subagents/preview",
|
||||
"POST /api/onboarding/complete",
|
||||
"GET /api/files/list",
|
||||
"GET /api/files/read",
|
||||
"GET /api/files/content",
|
||||
"GET /api/files/download",
|
||||
"POST /api/files/upload",
|
||||
"POST /api/files/mkdir",
|
||||
"POST /api/files/write",
|
||||
"POST /api/files/delete",
|
||||
"POST /api/files/transfer",
|
||||
"WS /ws",
|
||||
)
|
||||
|
|
@ -125,7 +125,9 @@ def _review_fields(
|
|||
github_token_configured: bool | None = None,
|
||||
) -> dict[str, Any]:
|
||||
stale = loaded.review.is_stale_for(loaded.content_hash) if stale is None else stale
|
||||
gate = skill_review_gate(loaded.review.status, stale=stale) if gate is None else gate
|
||||
gate = (skill_review_gate(loaded.review.status, stale=stale,
|
||||
findings=getattr(loaded.review, "findings", None))
|
||||
if gate is None else gate)
|
||||
source = str(getattr(loaded, "source", "") or "")
|
||||
official_hub_verified = False
|
||||
if source == "ouroboroshub":
|
||||
|
|
@ -387,7 +389,7 @@ def _build_extensions_index(drive_root, repo_path):
|
|||
}
|
||||
if bool(getattr(s, "identity_collision", False)):
|
||||
stale = True
|
||||
gate = skill_review_gate(s.review.status, stale=stale)
|
||||
gate = skill_review_gate(s.review.status, stale=stale, findings=s.review.findings)
|
||||
# Serialize the collision fact itself: hub_sync must fail closed
|
||||
# (no-action conflict card) instead of first-wins joining one of
|
||||
# several same-name occupants (scope-review reproduction).
|
||||
|
|
@ -776,7 +778,7 @@ async def api_skill_toggle(request: Request) -> JSONResponse:
|
|||
}
|
||||
stale = loaded.review.is_stale_for(loaded.content_hash)
|
||||
grants = grant_status_for_skill(drive_root, loaded)
|
||||
gate = skill_review_gate(loaded.review.status, stale=stale)
|
||||
gate = skill_review_gate(loaded.review.status, stale=stale, findings=loaded.review.findings)
|
||||
if not gate["executable_review"]:
|
||||
return {
|
||||
"error": "cannot enable until review status is a fresh executable review",
|
||||
|
|
@ -1003,7 +1005,8 @@ async def api_owner_skill_attest_review(request: Request) -> JSONResponse:
|
|||
log.debug("Failed to write owner attestation audit event", exc_info=True)
|
||||
if status != "clean":
|
||||
# Deterministic preflight floor failed, the skill is not owner-own, or it could not
|
||||
# be loaded/hashed: 409 — not attestable (existing review state is left untouched).
|
||||
# be loaded/hashed: 409 — not attestable. A failed preflight persists as the recorded
|
||||
# review result when review.json was absent/stale; a fresh valid verdict stays untouched.
|
||||
return JSONResponse(payload, status_code=409)
|
||||
return JSONResponse(payload)
|
||||
|
||||
|
|
@ -1245,7 +1248,7 @@ async def api_skill_grants(request: Request) -> JSONResponse:
|
|||
"status_code": 400,
|
||||
}
|
||||
stale = loaded.review.is_stale_for(loaded.content_hash)
|
||||
gate = skill_review_gate(loaded.review.status, stale=stale)
|
||||
gate = skill_review_gate(loaded.review.status, stale=stale, findings=loaded.review.findings)
|
||||
if not review_status_allows_execution(loaded.review.status) or stale:
|
||||
return {
|
||||
"error": "key and permission grants require a fresh executable review",
|
||||
|
|
|
|||
|
|
@ -23,7 +23,7 @@ from ouroboros.outcomes import normalize_outcome_axes
|
|||
from ouroboros.post_task_checkpoint import post_task_synthesis_is_open
|
||||
from ouroboros.subagent_messages import SUBAGENT_MESSAGE_FIELDS, subagent_message_meta
|
||||
from ouroboros.task_results import TASK_COST_META_FIELDS as _TASK_COST_META_FIELDS
|
||||
from ouroboros.utils import utc_now_iso
|
||||
from ouroboros.utils import strip_markdown, utc_now_iso
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
|
@ -259,7 +259,7 @@ def _user_annotation(
|
|||
key: annotation.get(key)
|
||||
for key in (
|
||||
"action", "target", "target_label", "status", "detail", "options",
|
||||
"attachment_manifest",
|
||||
"attachment_manifest", "routing_token",
|
||||
)
|
||||
if key in annotation
|
||||
}
|
||||
|
|
@ -440,6 +440,11 @@ def _chat_quota_predicate(row_matches_thread):
|
|||
return False
|
||||
if str(e.get("direction", "")).lower() not in ("in", "out"):
|
||||
return False
|
||||
if e.get("type") == "routing_options":
|
||||
# LLM-grounding rows (#198) are skipped by the render loop, so
|
||||
# they must not consume the human-row quota either — a run of
|
||||
# picker refusals would silently evict real messages.
|
||||
return False
|
||||
if is_a2a_chat_id(e.get("chat_id", 1)):
|
||||
return False
|
||||
ec = _stored_chat_id(e.get("chat_id"), 1)
|
||||
|
|
@ -798,6 +803,11 @@ def _collect_chat_rows(
|
|||
Returns ``(rows, quota_row_count)`` — the transformed history records and
|
||||
how many read entries satisfied the reader's quota predicate (feeds the
|
||||
window-truncation metadata)."""
|
||||
# Quiz lifecycle merge (#Q-2b): the chat row froze the card at ask time
|
||||
# ("open"); the durable truth lives in the owner_quiz task-result
|
||||
# projection. One projection read per distinct asking task, cached for
|
||||
# this call.
|
||||
quiz_projection_cache: Dict[str, Dict[str, Any]] = {}
|
||||
combined: list = []
|
||||
chat_quota_rows = 0
|
||||
stream_gaps: set[str] = set()
|
||||
|
|
@ -852,6 +862,10 @@ def _collect_chat_rows(
|
|||
# provenance surface.
|
||||
}
|
||||
if rec["system_type"] in {"project_started", "project_completion_summary"}:
|
||||
# Read-side plain normalization for lifecycle rows persisted
|
||||
# before the producer stripped markdown; a no-op on new rows.
|
||||
# The durable chat.jsonl is never rewritten.
|
||||
rec["text"] = strip_markdown(rec["text"])
|
||||
for key in ("project_id", "project_name", "target_label", "status"):
|
||||
if key in entry:
|
||||
rec[key] = str(entry.get(key) or "")
|
||||
|
|
@ -873,17 +887,40 @@ def _collect_chat_rows(
|
|||
# Delivered document rows carry lightweight media metadata (no
|
||||
# base64); surface a msg_type + download_url so the frontend
|
||||
# rebuilds the file bubble on reload instead of a bare text line.
|
||||
if entry.get("type") == "routing_options":
|
||||
# LLM-context grounding row (#198, decision 4=A): the web
|
||||
# surface renders the richer picker card from the annotation,
|
||||
# so the plain-text list never double-renders here.
|
||||
continue
|
||||
if entry.get("type") == "document":
|
||||
rec["msg_type"] = "document"
|
||||
rec["filename"] = str(entry.get("filename") or "file")
|
||||
rec["mime"] = str(entry.get("mime") or "application/octet-stream")
|
||||
rec["download_url"] = str(entry.get("download_url") or "")
|
||||
rec["caption"] = str(entry.get("caption") or "")
|
||||
if "size_bytes" in entry:
|
||||
rec["size_bytes"] = coerce_int(entry.get("size_bytes"), 0)
|
||||
elif entry.get("type") in {"photo", "video"} and entry.get("download_url"):
|
||||
rec["msg_type"] = str(entry["type"])
|
||||
rec["mime"] = str(entry.get("mime") or "")
|
||||
rec["download_url"] = str(entry["download_url"])
|
||||
rec["caption"] = str(entry.get("caption") or "")
|
||||
elif entry.get("type") == "links":
|
||||
rec.update(msg_type="links", actions=list(entry.get("actions") or []), title=str(entry.get("title") or ""))
|
||||
elif entry.get("type") == "quiz" and isinstance(entry.get("quiz"), dict):
|
||||
quiz = dict(entry["quiz"])
|
||||
_qtid = str(entry.get("task_id") or "")
|
||||
if _qtid:
|
||||
if _qtid not in quiz_projection_cache:
|
||||
from ouroboros.owner_quiz import quiz_states
|
||||
|
||||
quiz_projection_cache[_qtid] = quiz_states(chat_path.parent.parent, _qtid)
|
||||
_live = quiz_projection_cache[_qtid].get(str(quiz.get("quiz_id") or ""))
|
||||
if isinstance(_live, dict):
|
||||
quiz["state"] = str(_live.get("state") or quiz.get("state") or "open")
|
||||
if "answered_index" in _live:
|
||||
quiz["answered_index"] = _live["answered_index"]
|
||||
rec.update(msg_type="quiz", quiz=quiz)
|
||||
if "task_terminal_status" in entry:
|
||||
rec["task_terminal_status"] = str(entry.get("task_terminal_status") or "")
|
||||
_copy_task_summary_metadata(rec, entry)
|
||||
|
|
|
|||
|
|
@ -19,7 +19,7 @@ from starlette.responses import JSONResponse
|
|||
from starlette.routing import Route, WebSocketRoute
|
||||
from starlette.websockets import WebSocket, WebSocketDisconnect
|
||||
|
||||
from ouroboros.contracts.chat_id_policy import A2A_CHAT_ID_MAX, A2A_CHAT_ID_MIN
|
||||
from ouroboros.contracts.chat_id_policy import A2A_CHAT_ID_MAX, A2A_CHAT_ID_MIN, is_a2a_chat_id
|
||||
from ouroboros.event_bus import get_global_event_bus
|
||||
from ouroboros.skill_loader import (
|
||||
find_skill,
|
||||
|
|
@ -309,6 +309,16 @@ async def _api_chat_inject(request: Request) -> JSONResponse:
|
|||
bridge = ctx.bridge_getter()
|
||||
chat_id = int(payload.get("chat_id") or 0)
|
||||
wait_for_response = bool(payload.get("wait_for_response", False))
|
||||
if wait_for_response and not is_a2a_chat_id(chat_id):
|
||||
# A response subscription resolves on the FIRST non-progress frame
|
||||
# in the chat. On a human/project chat that frame can be any
|
||||
# concurrent task's answer — and, now that owner sends deliver
|
||||
# live mid-task, any proactive frame. Only A2A-allocated chats
|
||||
# (see /chat/allocate) have single-conversation semantics.
|
||||
return _json_error(
|
||||
"wait_for_response requires an A2A-allocated chat_id "
|
||||
"(allocate one via /chat/allocate-internal)", 400,
|
||||
)
|
||||
response_event: asyncio.Event = asyncio.Event()
|
||||
response_holder: dict[str, str] = {}
|
||||
if wait_for_response:
|
||||
|
|
|
|||
|
|
@ -568,7 +568,7 @@ def _installed_skill_payload(skill: Any, drive_root: pathlib.Path, *, provenance
|
|||
except Exception:
|
||||
payload_root = ""
|
||||
stale = skill.review.is_stale_for(skill.content_hash)
|
||||
gate = skill_review_gate(skill.review.status, stale=stale)
|
||||
gate = skill_review_gate(skill.review.status, stale=stale, findings=skill.review.findings)
|
||||
payload = {
|
||||
"name": skill.name,
|
||||
"type": skill.manifest.type,
|
||||
|
|
|
|||
|
|
@ -116,7 +116,8 @@ def collect_routes(
|
|||
from ouroboros.gateway.tasks import (
|
||||
api_task_artifact,
|
||||
api_task_cancel,
|
||||
api_task_hurry,
|
||||
api_decision_answer,
|
||||
api_task_hurry,
|
||||
api_task_resume,
|
||||
api_task_events,
|
||||
api_task_get,
|
||||
|
|
@ -238,6 +239,7 @@ def collect_routes(
|
|||
Route("/api/tasks/{task_id}/cancel", endpoint=api_task_cancel, methods=["POST"]),
|
||||
Route("/api/tasks/{task_id}/hurry", endpoint=api_task_hurry, methods=["POST"]),
|
||||
Route("/api/tasks/{task_id}/resume", endpoint=api_task_resume, methods=["POST"]),
|
||||
Route("/api/decisions", endpoint=api_decision_answer, methods=["POST"]),
|
||||
Route("/api/schedules", endpoint=api_schedules_list, methods=["GET"]),
|
||||
Route("/api/schedules", endpoint=api_schedules_upsert, methods=["POST"]),
|
||||
Route("/api/schedules/{schedule_id}", endpoint=api_schedules_delete, methods=["DELETE"]),
|
||||
|
|
|
|||
328
ouroboros/gateway/routing_decision.py
Normal file
328
ouroboros/gateway/routing_decision.py
Normal file
|
|
@ -0,0 +1,328 @@
|
|||
"""The routing family of ``POST /api/decisions`` — the #198 picker dispatcher.
|
||||
|
||||
An owner's click on a routing picker card executes the chosen action WITHOUT
|
||||
a new LLM turn, by reusing the exact supervisor handlers the LLM routing
|
||||
tools already use: the event goes into the same worker-event queue
|
||||
(``supervisor.workers.get_event_q``) that ``ouroboros/tools/control.py``
|
||||
feeds, and the outcome is read from the same durable receipts
|
||||
(``routing_wait`` — the task-result ``promotion_admission`` record and the
|
||||
chat-annotation receipt). No parallel steering/promotion machinery exists
|
||||
here; this module only validates the click against the durable refusal row
|
||||
and translates it into the established event shapes.
|
||||
|
||||
Idempotency without a new registry (owner decision 1=A): the dispatched
|
||||
event's ``routing_token`` and (for a new task) the ``task_id`` are DERIVED
|
||||
deterministically from ``(client_message_id, refusal routing_token,
|
||||
option_index)`` — a replayed request re-derives the same identities, so
|
||||
the supervisor's admission reservation and the steer mailbox ``msg_id``
|
||||
dedupe it instead of double-dispatching. After a confirmed dispatch the
|
||||
gateway appends one closing annotation row under the ORIGINAL refusal token
|
||||
carrying the ``request_id``, so replays read their own confirmation and a
|
||||
different click honestly loses with the card's true state.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import logging
|
||||
import pathlib
|
||||
from typing import Any, Dict, Tuple
|
||||
|
||||
from ouroboros.utils import utc_now_iso
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
_DISPATCH_STATUS_BY_ACTION = {"steer_task": "delivered", "new_task_in_project": "scheduled"}
|
||||
|
||||
|
||||
def parse_routing_decision_id(decision_id: str) -> Tuple[str, str]:
|
||||
"""``routing:{client_message_id}:{routing_token}`` → (cmid, token)."""
|
||||
parts = str(decision_id or "").split(":")
|
||||
if len(parts) < 3 or parts[0] != "routing" or not parts[1] or not parts[-1]:
|
||||
return "", ""
|
||||
return ":".join(parts[1:-1]), parts[-1]
|
||||
|
||||
|
||||
def _derived_identity(client_message_id: str, token: str, option_index: int) -> Tuple[str, str]:
|
||||
"""Deterministic (dispatch routing_token, task_id) for idempotent replays."""
|
||||
seed = f"{client_message_id}:{token}:{int(option_index)}".encode("utf-8")
|
||||
dispatch_token = hashlib.sha256(seed + b":token").hexdigest()
|
||||
task_id = hashlib.sha256(seed + b":task").hexdigest()[:16]
|
||||
return dispatch_token, task_id
|
||||
|
||||
|
||||
def _origin_message(drive_root: pathlib.Path, client_message_id: str) -> Dict[str, Any]:
|
||||
"""Recover the owner's ORIGINAL message row by client_message_id.
|
||||
|
||||
The picker transports the owner's exact ingress bytes to the chosen
|
||||
destination (the same rule the steer tool follows: a paraphrase must
|
||||
never replace the owner's text). Bounded reverse scan over the live
|
||||
chat log plus its rotation archives."""
|
||||
from ouroboros.gateway._helpers import read_rotated_jsonl_entries
|
||||
|
||||
target = str(client_message_id or "")
|
||||
|
||||
def _is_origin(entry: Dict[str, Any]) -> bool:
|
||||
return (
|
||||
str(entry.get("direction") or "") == "in"
|
||||
and str(entry.get("client_message_id") or "") == target
|
||||
)
|
||||
|
||||
try:
|
||||
entries = read_rotated_jsonl_entries(
|
||||
drive_root / "logs" / "chat.jsonl",
|
||||
drive_root / "archive",
|
||||
"chat",
|
||||
want=1,
|
||||
counts_toward_quota=_is_origin,
|
||||
)
|
||||
except Exception:
|
||||
log.debug("origin-message scan failed for %s", client_message_id, exc_info=True)
|
||||
return {}
|
||||
for entry in reversed(list(entries or [])):
|
||||
if isinstance(entry, dict) and _is_origin(entry):
|
||||
return entry
|
||||
return {}
|
||||
|
||||
|
||||
def handle_routing_decision(
|
||||
drive_root: pathlib.Path, *, request_id: str, decision_id: str,
|
||||
option_index: int, comment: str = "",
|
||||
) -> Tuple[int, Dict[str, Any]]:
|
||||
"""Execute one picker choice; returns (http_status, payload)."""
|
||||
client_message_id, token = parse_routing_decision_id(decision_id)
|
||||
if not client_message_id or not token:
|
||||
return 400, {"ok": False, "error": "malformed_decision_id",
|
||||
"decision_id": decision_id}
|
||||
from ouroboros.project_dialogue import (
|
||||
append_chat_annotation,
|
||||
chat_annotation_receipt,
|
||||
latest_chat_annotations,
|
||||
)
|
||||
|
||||
receipt = chat_annotation_receipt(drive_root, client_message_id, token)
|
||||
if not receipt:
|
||||
# The refusal row was superseded (a newer routing attempt re-minted
|
||||
# the token) or never existed — the card settles instead of retrying.
|
||||
latest = latest_chat_annotations(drive_root).get(client_message_id, {})
|
||||
return 409, {"ok": False, "error": "decision_superseded",
|
||||
"decision_id": decision_id,
|
||||
"state": "superseded",
|
||||
"latest_status": str(latest.get("status") or "")}
|
||||
status = str(receipt.get("status") or "")
|
||||
same_request = str(receipt.get("detail") or "") == f"request:{request_id}"
|
||||
if status == "dispatch_pending" and not same_request:
|
||||
# First-wins BEFORE the side effect (the quiz-card discipline): a
|
||||
# competing click while another request's dispatch is in flight is
|
||||
# refused; the winner's replay (same request_id) re-enters below.
|
||||
return 409, {"ok": False, "error": "dispatch_in_flight",
|
||||
"decision_id": decision_id, "state": "pending"}
|
||||
if status == "dispatch_pending" and same_request:
|
||||
# The replay must name the SAME option the claim bound: the dispatch
|
||||
# identity derives from the option, so a same-id different-option
|
||||
# request is a competing click wearing the winner's key.
|
||||
claimed_option = str(receipt.get("reason") or "")
|
||||
current_option = (
|
||||
f"claimed_option:{int(option_index)}"
|
||||
if isinstance(option_index, int) and not isinstance(option_index, bool)
|
||||
else ""
|
||||
)
|
||||
if claimed_option != current_option:
|
||||
return 409, {"ok": False, "error": "request_option_mismatch",
|
||||
"decision_id": decision_id, "state": "pending"}
|
||||
if status not in {"needs_manual_target", "dispatch_pending"}:
|
||||
if same_request:
|
||||
payload = {"ok": True, "decision_id": decision_id,
|
||||
"state": "answered", "duplicate": True,
|
||||
"dispatched": status}
|
||||
recorded = str(receipt.get("reason") or "")
|
||||
if recorded.startswith("answered_option:"):
|
||||
try:
|
||||
payload["answered_index"] = int(recorded.split(":", 1)[1])
|
||||
except ValueError:
|
||||
pass
|
||||
return 200, payload
|
||||
return 409, {"ok": False, "error": "decision_closed",
|
||||
"decision_id": decision_id, "state": "answered",
|
||||
"dispatched": status}
|
||||
options = receipt.get("options") if isinstance(receipt.get("options"), list) else []
|
||||
if not isinstance(option_index, int) or not (0 <= option_index < len(options)):
|
||||
return 400, {"ok": False, "error": "option_out_of_range",
|
||||
"decision_id": decision_id}
|
||||
chosen = options[option_index] if isinstance(options[option_index], dict) else {}
|
||||
action = str(chosen.get("action") or "")
|
||||
if action not in _DISPATCH_STATUS_BY_ACTION:
|
||||
return 400, {"ok": False, "error": "option_not_dispatchable",
|
||||
"decision_id": decision_id, "action": action}
|
||||
|
||||
origin = _origin_message(drive_root, client_message_id)
|
||||
# VERBATIM contract: the dispatched message is the owner's exact ingress
|
||||
# bytes; strip() is only the emptiness CHECK, never a rewrite.
|
||||
raw_origin_text = str(origin.get("text") or "")
|
||||
origin_text = raw_origin_text
|
||||
if not raw_origin_text.strip():
|
||||
# The card stays OPEN and truthful: nothing was dispatched, and the
|
||||
# durable refusal row still says needs_manual_target.
|
||||
return 409, {"ok": False, "error": "origin_text_unavailable",
|
||||
"decision_id": decision_id, "state": "open",
|
||||
"reason": "the original message could not be recovered"}
|
||||
if comment:
|
||||
origin_text = f"{origin_text}\n\n[Owner picker comment] {comment}"
|
||||
chat_id = int(origin.get("chat_id") or 0)
|
||||
|
||||
attachment_uploads = (
|
||||
[dict(row) for row in receipt.get("attachment_manifest") if isinstance(row, dict)]
|
||||
if isinstance(receipt.get("attachment_manifest"), list) else []
|
||||
)
|
||||
dispatch_token, derived_task_id = _derived_identity(client_message_id, token, option_index)
|
||||
# Origin provenance BY VALUE from the canonical row itself (the same rail
|
||||
# the LLM promote path rides): identity ref + full text + surface fact.
|
||||
from ouroboros.project_dialogue import build_owner_message_ref
|
||||
|
||||
provenance: Dict[str, Any] = {
|
||||
"source_ref": build_owner_message_ref(
|
||||
chat_id=chat_id, client_message_id=client_message_id,
|
||||
ts=str(origin.get("ts") or ""), text=raw_origin_text,
|
||||
),
|
||||
"source_text": raw_origin_text,
|
||||
}
|
||||
if isinstance(origin.get("client_surface"), dict) and origin.get("client_surface"):
|
||||
provenance["client_surface"] = dict(origin["client_surface"])
|
||||
if action == "steer_task":
|
||||
evt: Dict[str, Any] = {
|
||||
"type": "steer_task",
|
||||
"routing_token": dispatch_token,
|
||||
"target_task_id": str(chosen.get("task_id") or ""),
|
||||
"message": origin_text,
|
||||
"chat_id": chat_id,
|
||||
"client_message_id": client_message_id,
|
||||
# The option list was host-built for this exact message's lane and
|
||||
# the owner picked the row explicitly — global root addressing is
|
||||
# the validated intent, not a widening.
|
||||
"allow_global_root": True,
|
||||
"attachment_uploads": attachment_uploads,
|
||||
**provenance,
|
||||
"ts": utc_now_iso(),
|
||||
}
|
||||
else:
|
||||
evt = {
|
||||
"type": "promote_chat_to_task",
|
||||
"task_id": derived_task_id,
|
||||
"routing_token": dispatch_token,
|
||||
"objective": origin_text,
|
||||
"project_id": str(chosen.get("project_id") or ""),
|
||||
"chat_id": chat_id,
|
||||
# Label-only provenance for the receipt: the picker click IS a
|
||||
# routing decision on a refused message, so it wears the
|
||||
# route_to_project receipt label regardless of source chat.
|
||||
"routed_from_main": True,
|
||||
"client_message_id": client_message_id,
|
||||
"attachment_uploads": attachment_uploads,
|
||||
**provenance,
|
||||
"ts": utc_now_iso(),
|
||||
}
|
||||
|
||||
def _reopen_refusal() -> None:
|
||||
"""Restore the actionable refusal as the LATEST row (same token)."""
|
||||
try:
|
||||
# Token-bound: if a NEWER routing attempt re-minted the card while
|
||||
# our dispatch settled, its authority wins and this reopen is a
|
||||
# no-op (the allowed set covers only our claim and our dispatch
|
||||
# receipt, never a foreign token).
|
||||
append_chat_annotation(
|
||||
drive_root, client_message_id, action="route_decision",
|
||||
status="needs_manual_target", routing_token=token,
|
||||
options=options,
|
||||
attachment_manifest=receipt.get("attachment_manifest"),
|
||||
require_latest_token={token, dispatch_token},
|
||||
)
|
||||
except Exception:
|
||||
log.debug("refusal reopen append failed", exc_info=True)
|
||||
|
||||
# Claim BEFORE the side effect (first-wins): a compare-and-append under
|
||||
# the annotations lock — two concurrent clicks cannot both claim. The
|
||||
# claim row carries the options so a crash/reload before settlement still
|
||||
# renders honestly and the winner's replay re-enters with full validation
|
||||
# authority (its own claim is already latest, so it skips this step).
|
||||
if status == "needs_manual_target":
|
||||
claimed = append_chat_annotation(
|
||||
drive_root, client_message_id, action="route_decision",
|
||||
status="dispatch_pending", routing_token=token,
|
||||
detail=f"request:{request_id}", options=options,
|
||||
reason=f"claimed_option:{int(option_index)}",
|
||||
attachment_manifest=receipt.get("attachment_manifest"),
|
||||
require_latest_status={"needs_manual_target"},
|
||||
require_latest_token={token},
|
||||
)
|
||||
if not claimed:
|
||||
# Lost the claim race (or a transient lock failure) — the card
|
||||
# reads pending; a later click sees the settled truth.
|
||||
return 409, {"ok": False, "error": "dispatch_in_flight",
|
||||
"decision_id": decision_id, "state": "pending"}
|
||||
try:
|
||||
from multiprocessing.reduction import ForkingPickler
|
||||
|
||||
ForkingPickler.dumps(dict(evt))
|
||||
from supervisor.workers import get_event_q
|
||||
|
||||
get_event_q().put_nowait(dict(evt))
|
||||
except Exception as exc:
|
||||
log.warning("routing decision dispatch failed", exc_info=True)
|
||||
_reopen_refusal() # nothing is in flight — the card stays clickable
|
||||
return 503, {"ok": False, "error": "dispatch_unavailable",
|
||||
"decision_id": decision_id,
|
||||
"detail": f"{type(exc).__name__}: {exc}"}
|
||||
|
||||
from ouroboros.routing_wait import (
|
||||
wait_for_promotion_admission,
|
||||
wait_for_routing_annotation,
|
||||
)
|
||||
|
||||
if action == "steer_task":
|
||||
outcome = wait_for_routing_annotation(drive_root, client_message_id, dispatch_token)
|
||||
else:
|
||||
outcome = wait_for_promotion_admission(
|
||||
drive_root, derived_task_id, dispatch_token,
|
||||
client_message_id=client_message_id,
|
||||
)
|
||||
outcome_status = str(outcome.get("status") or "unconfirmed")
|
||||
expected = _DISPATCH_STATUS_BY_ACTION[action]
|
||||
if outcome_status == expected:
|
||||
try:
|
||||
# The closing row IS the presentation receipt after hydration, so
|
||||
# it wears the real routing action + label ("Steered task · X"),
|
||||
# plus the request id the replay path reads back.
|
||||
append_chat_annotation(
|
||||
drive_root, client_message_id,
|
||||
action="steer_task" if action == "steer_task" else "promote_chat_to_task",
|
||||
target=str(chosen.get("task_id") or chosen.get("project_id") or ""),
|
||||
target_label=str(chosen.get("label") or chosen.get("title")
|
||||
or chosen.get("project_name") or ""),
|
||||
status=expected, routing_token=token,
|
||||
detail=f"request:{request_id}",
|
||||
reason=f"answered_option:{int(option_index)}",
|
||||
)
|
||||
except Exception:
|
||||
log.debug("closing annotation append failed", exc_info=True)
|
||||
payload: Dict[str, Any] = {"ok": True, "decision_id": decision_id,
|
||||
"state": "answered", "dispatched": expected,
|
||||
"answered_index": int(option_index)}
|
||||
if action != "steer_task":
|
||||
payload["task_id"] = derived_task_id
|
||||
return 200, payload
|
||||
if outcome_status in {"rejected", "needs_manual_target"}:
|
||||
# The handler's rejection receipt (under the DISPATCH token) is now
|
||||
# the latest row — re-assert the refusal under the ORIGINAL token so
|
||||
# the card the UI re-opens still validates and replays cleanly.
|
||||
_reopen_refusal()
|
||||
return 409, {"ok": False, "error": "dispatch_rejected",
|
||||
"decision_id": decision_id, "state": "open",
|
||||
"reason": str(outcome.get("reason") or outcome_status)}
|
||||
# Unconfirmed: honestly retriable — the derived identities make a replay
|
||||
# of the SAME request byte-identical, so the supervisor dedupes it.
|
||||
return 503, {"ok": False, "error": "dispatch_unconfirmed",
|
||||
"decision_id": decision_id,
|
||||
"reason": str(outcome.get("reason") or "confirmation_timeout")}
|
||||
|
||||
|
||||
__all__ = ["handle_routing_decision", "parse_routing_decision_id"]
|
||||
|
|
@ -181,32 +181,47 @@ def _rehydrate_mcp_servers_payload(incoming: Any, current: Any) -> list:
|
|||
|
||||
_IMMEDIATE_KEYS = frozenset({
|
||||
"TOTAL_BUDGET",
|
||||
"OUROBOROS_SOFT_TIMEOUT_SEC",
|
||||
"OUROBOROS_HARD_TIMEOUT_SEC",
|
||||
# The OUTER per-call tool cap reads settings.json BEFORE env on every tool
|
||||
# call in every process (loop_tool_execution.py), so a saved change bites
|
||||
# the currently running task's next tool call. The inner shell subprocess
|
||||
# timeout still prefers the worker env (next task) — disclosed residual.
|
||||
"OUROBOROS_TOOL_TIMEOUT_SEC",
|
||||
"GITHUB_TOKEN",
|
||||
"GITHUB_REPO",
|
||||
"OUROBOROS_UPDATE_CHANNEL",
|
||||
# The save handler hot-reconfigures MCP itself before responding
|
||||
# (_apply_settings_save_side_effects), and worker processes re-check the
|
||||
# settings mtime on their next tool-schema read; a reconfigure failure is
|
||||
# surfaced as a save warning instead of silently keeping the claim.
|
||||
"MCP_ENABLED",
|
||||
"MCP_SERVERS",
|
||||
"MCP_TOOL_TIMEOUT_SEC",
|
||||
})
|
||||
|
||||
# Retired keys the settings document still carries: changing them has NO
|
||||
# effect at any point (supervisor/queue.py accepts them as typed no-ops).
|
||||
# Classifying them as immediate or next-task would both be lies; the save
|
||||
# reports them with an explicit "retired" warning instead.
|
||||
_RETIRED_NO_EFFECT_KEYS = frozenset({
|
||||
"OUROBOROS_SOFT_TIMEOUT_SEC",
|
||||
"OUROBOROS_HARD_TIMEOUT_SEC",
|
||||
})
|
||||
|
||||
_RESTART_REQUIRED_KEYS = frozenset({
|
||||
"OUROBOROS_MAX_WORKERS",
|
||||
"OUROBOROS_SERVER_HOST",
|
||||
# The host-service port is bound once at server startup.
|
||||
"OUROBOROS_HOST_SERVICE_PORT",
|
||||
# Pooled workers load the extension registry once at spawn and never
|
||||
# reload it per task; the save-time server reload keeps the skills UI
|
||||
# fresh, but agent tasks see the new repo only after a restart.
|
||||
"OUROBOROS_SKILLS_REPO_PATH",
|
||||
"LOCAL_MODEL_SOURCE",
|
||||
"LOCAL_MODEL_FILENAME",
|
||||
"LOCAL_MODEL_PORT",
|
||||
"LOCAL_MODEL_N_GPU_LAYERS",
|
||||
"LOCAL_MODEL_CONTEXT_LENGTH",
|
||||
"LOCAL_MODEL_CHAT_FORMAT",
|
||||
"OPENAI_BASE_URL",
|
||||
"OPENAI_COMPATIBLE_BASE_URL",
|
||||
"CLOUDRU_FOUNDATION_MODELS_BASE_URL",
|
||||
# Region selects the MiniMax base URL (api.minimax.io vs api.minimaxi.com),
|
||||
# so it changes routing exactly like the base-URL keys above it.
|
||||
"MINIMAX_REGION",
|
||||
"GIGACHAT_SCOPE",
|
||||
"GIGACHAT_BASE_URL",
|
||||
"GIGACHAT_VERIFY_SSL_CERTS",
|
||||
# Background cognition reads these at consciousness __init__, so a change
|
||||
# only takes effect after restart (Phase 4 Evolution settings group).
|
||||
"OUROBOROS_BG_WAKEUP_MIN",
|
||||
|
|
@ -226,6 +241,30 @@ def _classify_settings_changes(
|
|||
]
|
||||
|
||||
|
||||
def _effect_buckets(all_changed: list, warnings: list) -> tuple:
|
||||
"""Split changed keys into the honest effect buckets for the save response.
|
||||
|
||||
A retired key applies NEVER, so neither bucket may claim it — it is
|
||||
reported through an explicit warning instead (#285).
|
||||
"""
|
||||
retired_changed = sorted(k for k in all_changed if k in _RETIRED_NO_EFFECT_KEYS)
|
||||
if retired_changed:
|
||||
warnings.append(
|
||||
"Retired setting(s) saved, but they no longer affect anything: "
|
||||
+ ", ".join(retired_changed)
|
||||
+ ". Task runtime is governed by OUROBOROS_TASK_IDLE_TIMEOUT_SEC "
|
||||
"and OUROBOROS_TASK_ABS_CEILING_SEC."
|
||||
)
|
||||
immediate_changed = [k for k in all_changed if k in _IMMEDIATE_KEYS]
|
||||
next_task_changed = [
|
||||
k for k in all_changed
|
||||
if k not in _IMMEDIATE_KEYS
|
||||
and k not in _RESTART_REQUIRED_KEYS
|
||||
and k not in _RETIRED_NO_EFFECT_KEYS
|
||||
]
|
||||
return immediate_changed, next_task_changed
|
||||
|
||||
|
||||
def _merge_settings_payload(current: Dict[str, Any], body: Dict[str, Any]) -> Dict[str, Any]:
|
||||
merged = {k: v for k, v in current.items()}
|
||||
for key in _SETTINGS_DEFAULTS:
|
||||
|
|
@ -940,8 +979,14 @@ def _apply_settings_save_side_effects(
|
|||
current: Dict[str, Any],
|
||||
old_effective_settings: Dict[str, Any],
|
||||
all_changed: list,
|
||||
) -> None:
|
||||
"""Post-save hot-reload side effects (MCP, extensions, supervisor budgets/timeouts)."""
|
||||
) -> list:
|
||||
"""Post-save hot-reload side effects (MCP, extensions, supervisor budgets/timeouts).
|
||||
|
||||
Returns warning strings for side effects that FAILED: an immediate-classed
|
||||
key whose hot apply broke must not let the save report "took effect
|
||||
immediately" without saying so (#285).
|
||||
"""
|
||||
side_effect_warnings: list = []
|
||||
if any(k in all_changed for k in ("MCP_ENABLED", "MCP_SERVERS", "MCP_TOOL_TIMEOUT_SEC")):
|
||||
try:
|
||||
from ouroboros.mcp_client import (
|
||||
|
|
@ -950,8 +995,13 @@ def _apply_settings_save_side_effects(
|
|||
)
|
||||
_mcp_reconfigure(current)
|
||||
_mcp_refresh_background(reason="settings")
|
||||
except Exception:
|
||||
except Exception as exc:
|
||||
log.warning("MCP reconfigure after settings change failed", exc_info=True)
|
||||
side_effect_warnings.append(
|
||||
"MCP reconfigure failed in the server process: "
|
||||
f"{type(exc).__name__}: {exc}. The saved values are on disk; "
|
||||
"agent processes retry on their next tool-schema read."
|
||||
)
|
||||
|
||||
# Skills repo/runtime changes require extension loader reconciliation.
|
||||
try:
|
||||
|
|
@ -980,8 +1030,13 @@ def _apply_settings_save_side_effects(
|
|||
_load_settings,
|
||||
repo_path=new_path or None,
|
||||
)
|
||||
except Exception:
|
||||
except Exception as exc:
|
||||
log.error("Extension reload after settings change failed", exc_info=True)
|
||||
side_effect_warnings.append(
|
||||
"Skills repo reload failed in the server process: "
|
||||
f"{type(exc).__name__}: {exc}. The saved path is on disk and "
|
||||
"applies after a restart."
|
||||
)
|
||||
|
||||
try:
|
||||
from supervisor.state import refresh_budget_from_settings
|
||||
|
|
@ -1000,6 +1055,7 @@ def _apply_settings_save_side_effects(
|
|||
refresh_budget_limit(new_budget)
|
||||
except Exception:
|
||||
pass
|
||||
return side_effect_warnings
|
||||
|
||||
|
||||
async def api_settings_post(request: Request) -> JSONResponse:
|
||||
|
|
@ -1238,10 +1294,12 @@ def _api_settings_post_locked(request: Request, body: Any) -> JSONResponse:
|
|||
_start_supervisor_if_needed_for_request(request, current)
|
||||
|
||||
boundary.at("hot-reload")
|
||||
_apply_settings_save_side_effects(request, current, old_effective_settings, all_changed)
|
||||
side_effect_warnings = _apply_settings_save_side_effects(
|
||||
request, current, old_effective_settings, all_changed)
|
||||
boundary.at("post-save notices")
|
||||
|
||||
warnings = []
|
||||
# Tolerate stubbed side effects returning None (test harnesses).
|
||||
warnings = list(side_effect_warnings or [])
|
||||
if _reviewer_fallback_warning:
|
||||
warnings.append(_reviewer_fallback_warning)
|
||||
if provider_defaults_changed:
|
||||
|
|
@ -1292,11 +1350,7 @@ def _api_settings_post_locked(request: Request, body: Any) -> JSONResponse:
|
|||
settings_to_save["GITHUB_REPO"] = resolved_slug
|
||||
_owner_write_settings(settings_to_save)
|
||||
os.environ["GITHUB_REPO"] = resolved_slug
|
||||
immediate_changed = [k for k in all_changed if k in _IMMEDIATE_KEYS]
|
||||
next_task_changed = [
|
||||
k for k in all_changed
|
||||
if k not in _IMMEDIATE_KEYS and k not in _RESTART_REQUIRED_KEYS
|
||||
]
|
||||
immediate_changed, next_task_changed = _effect_buckets(all_changed, warnings)
|
||||
agent_task_running = bool(next_task_changed) and started_before_save
|
||||
if agent_task_running:
|
||||
# Owner decision (2026-08-05): the task-start snapshot boundary STAYS —
|
||||
|
|
|
|||
|
|
@ -234,6 +234,11 @@ def _chat_activities_snapshot_safe(drive_root: Any, task_bindings: Any = None) -
|
|||
bindings = task_bindings if isinstance(task_bindings, dict) else {}
|
||||
with queue_mod._queue_lock:
|
||||
pending_rows = [dict(task) for task in queue_mod.PENDING]
|
||||
fence_rows = {
|
||||
str(key): dict(value)
|
||||
for key, value in queue_mod.BUDGET_ROOT_FENCES.items()
|
||||
if isinstance(value, dict)
|
||||
}
|
||||
running_rows = [
|
||||
(
|
||||
str(task_id),
|
||||
|
|
@ -269,10 +274,15 @@ def _chat_activities_snapshot_safe(drive_root: Any, task_bindings: Any = None) -
|
|||
"started_at": started_at,
|
||||
}
|
||||
|
||||
from supervisor.queue_transitions import budget_pause_fact
|
||||
|
||||
for row in pending_rows:
|
||||
task_id = str(row.get("id") or "")
|
||||
if task_id and _is_root(task_id, row):
|
||||
activities.append(_activity(task_id, row, "queued", _epoch_or_zero(row.get("queued_at"))))
|
||||
# #322 (P1): a budget-paused member must not masquerade as
|
||||
# "queued" — nothing will dispatch it until an explicit resume.
|
||||
phase = "budget_paused" if budget_pause_fact(row, fence_rows) else "queued"
|
||||
activities.append(_activity(task_id, row, phase, _epoch_or_zero(row.get("queued_at"))))
|
||||
for task_id, row, started_at in running_rows:
|
||||
if task_id and _is_root(task_id, row):
|
||||
phase = "finalizing" if _managed_task_finalizing(drive_root, task_id) else "working"
|
||||
|
|
|
|||
281
ouroboros/gateway/task_decision.py
Normal file
281
ouroboros/gateway/task_decision.py
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
"""``POST /api/decisions`` — the ONE owner decision-card answer ingress.
|
||||
|
||||
Owner decision 1=A: one UI component (``chat_decision.js``) and one incoming
|
||||
answer contract with an idempotent ``request_id``; ``decision_id`` composes
|
||||
EXISTING identities per family instead of minting a durable registry:
|
||||
|
||||
- ``quiz:{task_id}:{quiz_id}`` — served here (#Q-2b);
|
||||
- ``routing:{client_message_id}:{routing_token}`` — the #198 picker family,
|
||||
dispatched to ``gateway/routing_decision.py``;
|
||||
- ``interaction:{task_id}:{run_id}:{interaction_id}`` — RESERVED. #204 is
|
||||
served by the escalation hierarchy instead (owner decision 31): a delegated
|
||||
run's question wakes its nanny, who answers from task context via
|
||||
delegate_answer or escalates upward with the escalate verb — the owner sees
|
||||
a quiz card only when no ancestor answers, so no direct interaction card
|
||||
exists and this family stays a typed 501.
|
||||
|
||||
The quiz path mirrors the hurry ingress split (``gateway/task_hurry.py``):
|
||||
projection write first (request-id idempotent, first answer wins), then the
|
||||
typed ``KIND_QUIZ_ANSWER`` mailbox control on the task's physical drive, then
|
||||
the live ``quiz_state`` broadcast. A late answer to a settled task is an
|
||||
honest 409 carrying the card's true lifecycle state — the card settles
|
||||
instead of inviting retries.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from typing import Any, Dict, Optional, Tuple
|
||||
|
||||
from starlette.requests import Request
|
||||
from starlette.responses import JSONResponse
|
||||
|
||||
from ouroboros.gateway._helpers import json_error, json_exception, request_drive_root, request_json_or
|
||||
from ouroboros.task_results import resolve_task_lineage, validate_task_id
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
_REQUEST_ID_MAX = 128
|
||||
_COMMENT_MAX = 2000
|
||||
_SERVED_FAMILIES = {"quiz", "routing"}
|
||||
_KNOWN_FAMILIES = {"quiz", "routing", "interaction"}
|
||||
|
||||
|
||||
def _parse_quiz_decision_id(decision_id: str) -> Tuple[str, str, str]:
|
||||
"""Split ``quiz:{task_id}:{quiz_id}`` → (family, task_id, quiz_id)."""
|
||||
parts = str(decision_id or "").split(":", 2)
|
||||
family = parts[0] if parts else ""
|
||||
if family != "quiz" or len(parts) != 3 or not parts[1] or not parts[2]:
|
||||
return family, "", ""
|
||||
return family, parts[1], parts[2]
|
||||
|
||||
|
||||
def _live_root_task(task_id: str) -> Tuple[Optional[Dict[str, Any]], str]:
|
||||
"""Queue-lock read: the live task row, or a refusal reason.
|
||||
|
||||
``task_not_live`` is NOT a hard refusal for a quiz answer — the caller
|
||||
consults the durable projection for the honest late-answer state."""
|
||||
from supervisor import queue as q
|
||||
|
||||
with q._queue_lock:
|
||||
task: Optional[Dict[str, Any]] = None
|
||||
meta = q.RUNNING.get(task_id) if isinstance(q.RUNNING, dict) else None
|
||||
if isinstance(meta, dict) and isinstance(meta.get("task"), dict):
|
||||
task = dict(meta["task"])
|
||||
if task is None:
|
||||
task = next(
|
||||
(
|
||||
dict(row) for row in list(q.PENDING or [])
|
||||
if isinstance(row, dict) and str(row.get("id") or "") == task_id
|
||||
),
|
||||
None,
|
||||
)
|
||||
if task is None:
|
||||
return None, "task_not_live"
|
||||
lineage = resolve_task_lineage(
|
||||
task_id,
|
||||
metadata=task.get("metadata"),
|
||||
root_task_id=task.get("root_task_id"),
|
||||
parent_task_id=task.get("parent_task_id"),
|
||||
delegation_role=task.get("delegation_role"),
|
||||
original_task_id=task.get("original_task_id"),
|
||||
timeout_retry_from=task.get("timeout_retry_from"),
|
||||
)
|
||||
if not bool(lineage["is_root_task"]):
|
||||
# Decision-31 hierarchy: owner quiz cards come only from ROOT
|
||||
# tasks (a subagent escalates to its parent, never to a card).
|
||||
return None, "not_a_root_task"
|
||||
return task, ""
|
||||
|
||||
|
||||
def _quiz_answer_frame(block: Dict[str, Any], option_index: int, comment: str) -> str:
|
||||
"""Host-authored structural frame around the owner's VERBATIM choice.
|
||||
|
||||
The asked/answered timestamps ride inside so the MODEL judges freshness
|
||||
itself (owner decision 30=A — no host staleness verdict)."""
|
||||
options = block.get("options") if isinstance(block.get("options"), list) else []
|
||||
label = str(options[option_index]) if 0 <= option_index < len(options) else ""
|
||||
lines = [
|
||||
f"[Owner quiz answer] quiz {block.get('quiz_id')} — asked {block.get('asked_at')}, "
|
||||
f"answered {block.get('answered_at')}.",
|
||||
f"Question was: {block.get('question')}",
|
||||
f"The owner chose option {option_index + 1}: {label}",
|
||||
]
|
||||
if comment:
|
||||
lines.append(f"Owner comment (verbatim): {comment}")
|
||||
if str(block.get("assumption") or ""):
|
||||
lines.append(
|
||||
f"You continued under the assumption: {block.get('assumption')} — "
|
||||
"judge yourself whether work has moved past the answered fork."
|
||||
)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
async def api_decision_answer(request: Request) -> JSONResponse:
|
||||
"""POST /api/decisions — idempotent owner answer for a decision card."""
|
||||
body = await request_json_or(request, {})
|
||||
if not isinstance(body, dict):
|
||||
return json_error("request body must be a JSON object", 400)
|
||||
request_id = str(body.get("request_id") or "").strip()
|
||||
if not request_id or len(request_id) > _REQUEST_ID_MAX:
|
||||
return json_error(
|
||||
"request_id is required (a stable client-generated id, reused on retry)",
|
||||
400, reason_code="request_id_required",
|
||||
)
|
||||
decision_id = str(body.get("decision_id") or "").strip()
|
||||
raw_comment = body.get("comment")
|
||||
if raw_comment is not None and not isinstance(raw_comment, str):
|
||||
return json_error("comment must be a string", 400, reason_code="comment_invalid")
|
||||
comment = (raw_comment or "").strip()
|
||||
if len(comment) > _COMMENT_MAX:
|
||||
# VERBATIM contract: the frame signs the comment as the owner's exact
|
||||
# words — refuse instead of silently truncating them.
|
||||
return json_error(
|
||||
f"comment is {len(comment):,} characters (limit {_COMMENT_MAX:,}) — "
|
||||
"shorten it; it is delivered verbatim",
|
||||
400, reason_code="comment_too_long",
|
||||
)
|
||||
if any(key not in {"request_id", "decision_id", "option_index", "comment"} for key in body):
|
||||
return json_error(
|
||||
"decision accepts only {request_id, decision_id, option_index, comment?}",
|
||||
400, reason_code="unexpected_fields",
|
||||
)
|
||||
family, task_id, quiz_id = _parse_quiz_decision_id(decision_id)
|
||||
if family not in _KNOWN_FAMILIES:
|
||||
return json_error(
|
||||
"unknown decision family (expected quiz:/routing:/interaction:)",
|
||||
400, reason_code="unknown_decision_family",
|
||||
)
|
||||
if family not in _SERVED_FAMILIES:
|
||||
# Typed, honest: the interaction family is RESERVED — #204 is served
|
||||
# by the escalation hierarchy (see the module docstring), so no direct
|
||||
# owner interaction card exists by design.
|
||||
return json_error(
|
||||
f"the {family} decision family is not served yet",
|
||||
501, reason_code="decision_family_not_served",
|
||||
)
|
||||
raw_index = body.get("option_index")
|
||||
if not isinstance(raw_index, int) or isinstance(raw_index, bool) or raw_index < 0:
|
||||
return json_error(
|
||||
"option_index must be a non-negative integer",
|
||||
400, reason_code="option_index_invalid",
|
||||
)
|
||||
drive_root = request_drive_root(request)
|
||||
if family == "routing":
|
||||
from ouroboros.gateway.routing_decision import handle_routing_decision
|
||||
|
||||
status_code, payload = await asyncio.to_thread(
|
||||
handle_routing_decision, drive_root,
|
||||
request_id=request_id, decision_id=decision_id,
|
||||
option_index=raw_index, comment=comment,
|
||||
)
|
||||
return JSONResponse(payload, status_code=status_code)
|
||||
if not task_id or not quiz_id:
|
||||
return json_error(
|
||||
"malformed quiz decision_id (expected quiz:{task_id}:{quiz_id})",
|
||||
400, reason_code="malformed_decision_id",
|
||||
)
|
||||
try:
|
||||
task_id = validate_task_id(task_id)
|
||||
except ValueError as exc:
|
||||
return json_error(str(exc), 400)
|
||||
try:
|
||||
task, refusal = _live_root_task(task_id)
|
||||
if refusal == "not_a_root_task":
|
||||
return json_error(
|
||||
"quiz answers address root tasks only", 409,
|
||||
task_id=task_id, reason_code=refusal,
|
||||
)
|
||||
from ouroboros.owner_quiz import record_answered, reconcile_terminal
|
||||
|
||||
if task is None:
|
||||
# The author is gone. Normally the task-done seam already expired
|
||||
# its open quizzes; a crash window can leave one open — heal it
|
||||
# here so a late answer gets the honest expired 409 instead of a
|
||||
# 200 recorded into a mailbox nobody will ever drain.
|
||||
reconcile_terminal(drive_root, task_id)
|
||||
outcome = record_answered(
|
||||
drive_root, task_id,
|
||||
quiz_id=quiz_id, option_index=raw_index,
|
||||
request_id=request_id, comment=comment,
|
||||
)
|
||||
if not outcome.get("ok"):
|
||||
error = str(outcome.get("error") or "quiz_answer_refused")
|
||||
state = str(outcome.get("state") or "")
|
||||
if error == "quiz_not_found":
|
||||
return json_error("quiz not found", 404, task_id=task_id,
|
||||
reason_code=error)
|
||||
status = 409
|
||||
payload: Dict[str, Any] = {
|
||||
"ok": False, "error": error, "decision_id": decision_id,
|
||||
}
|
||||
# The truthful lifecycle state settles the card client-side: a
|
||||
# closed quiz on a SETTLED task reads as expired, an already
|
||||
# answered one as answered.
|
||||
payload["state"] = state or ("expired_terminal" if task is None else "")
|
||||
refused_block = outcome.get("block") if isinstance(outcome.get("block"), dict) else {}
|
||||
if isinstance(refused_block.get("answered_index"), int):
|
||||
# The loser of a first-wins race settles honestly: the card
|
||||
# learns the WINNING option, never a false expiry.
|
||||
payload["answered_index"] = refused_block["answered_index"]
|
||||
if error == "option_out_of_range":
|
||||
status = 400
|
||||
return JSONResponse(payload, status_code=status)
|
||||
block = outcome.get("block") if isinstance(outcome.get("block"), dict) else {}
|
||||
if task is not None:
|
||||
from supervisor.queue import _task_drive_for_task
|
||||
|
||||
from ouroboros.owner_mailbox import KIND_QUIZ_ANSWER, write_owner_message
|
||||
|
||||
# EVERY accepted request appends the control — fresh, same-id
|
||||
# retry, or a duplicate after a mailbox write failure (the hurry
|
||||
# heal semantics): the msg_id is stable per quiz, so the drain
|
||||
# dedupes a doubled line while a LOST control is healed by any
|
||||
# retry instead of being unrecoverable (the drain reads only the
|
||||
# mailbox, never the projection).
|
||||
# TOCTOU residual (disclosed): the task can settle between the
|
||||
# liveness read and this append — the control then waits for a
|
||||
# same-id retry attempt (reset_attempt_controls_for_retry revokes
|
||||
# only hurry/finalize kinds), and the model judges freshness from
|
||||
# the frame's stamps (30=A). No lock spans both writes on purpose.
|
||||
answered_index = (int(block["answered_index"])
|
||||
if isinstance(block.get("answered_index"), int)
|
||||
else raw_index)
|
||||
frame = _quiz_answer_frame(block, answered_index, str(block.get("comment") or comment))
|
||||
drive = _task_drive_for_task(task, task_id)
|
||||
if not write_owner_message(
|
||||
drive, frame, task_id,
|
||||
msg_id=f"quiz_answer:{quiz_id}", kind=KIND_QUIZ_ANSWER,
|
||||
):
|
||||
# The projection already recorded the answer (the card is
|
||||
# truthful); the injection control failed — say so. A retry
|
||||
# of this request re-attempts the append.
|
||||
return json_error(
|
||||
"the answer was recorded but the task control could not "
|
||||
"be written — retry to deliver it to the task",
|
||||
503, task_id=task_id, reason_code="mailbox_write_failed",
|
||||
)
|
||||
try:
|
||||
from supervisor.message_bus import get_bridge
|
||||
|
||||
get_bridge().send_quiz_state(
|
||||
quiz_id, task_id, str(outcome.get("state") or "answered"),
|
||||
answered_index=block.get("answered_index"),
|
||||
)
|
||||
except Exception:
|
||||
log.debug("quiz_state broadcast failed for %s", quiz_id, exc_info=True)
|
||||
except Exception as exc:
|
||||
return json_exception(exc, 503)
|
||||
return JSONResponse({
|
||||
"ok": True,
|
||||
"decision_id": decision_id,
|
||||
"state": str(outcome.get("state") or "answered"),
|
||||
"answered_index": (int(block["answered_index"])
|
||||
if isinstance(block.get("answered_index"), int)
|
||||
else raw_index),
|
||||
"duplicate": bool(outcome.get("duplicate")),
|
||||
})
|
||||
|
||||
|
||||
__all__ = ["api_decision_answer"]
|
||||
|
|
@ -33,6 +33,7 @@ from ouroboros.gateway.task_events import ( # noqa: F401
|
|||
# Re-exported hurry ingress (same module-size split as task_events): route
|
||||
# wiring and tests address gateway.tasks.api_task_hurry.
|
||||
from ouroboros.gateway.task_hurry import api_task_hurry # noqa: F401
|
||||
from ouroboros.gateway.task_decision import api_decision_answer # noqa: F401
|
||||
from ouroboros.headless import (
|
||||
ARTIFACTS_DIR,
|
||||
ARTIFACT_STATUS_FAILED,
|
||||
|
|
@ -1552,6 +1553,7 @@ def _supervisor_ready_error(request: Request) -> Optional[JSONResponse]:
|
|||
__all__ = [
|
||||
"api_task_artifact",
|
||||
"api_task_cancel",
|
||||
"api_decision_answer",
|
||||
"api_task_hurry",
|
||||
"api_task_resume",
|
||||
"api_task_events",
|
||||
|
|
|
|||
|
|
@ -173,6 +173,9 @@ def _attempt_request(
|
|||
prompt_chars = len(json.dumps(prompt_payload, ensure_ascii=False, default=str))
|
||||
except Exception:
|
||||
prompt_chars = len(str(prompt_payload or ""))
|
||||
from ouroboros.context_fit import bounded_prompt_tokens_for_payload
|
||||
|
||||
bounded_tokens = bounded_prompt_tokens_for_payload(prompt_payload, prompt_chars)
|
||||
request_source = source
|
||||
if request_source is None:
|
||||
bound_scope = current_usage_scope()
|
||||
|
|
@ -199,6 +202,7 @@ def _attempt_request(
|
|||
candidate_measurement_kind="canonical_json_v1",
|
||||
physical_context=current_physical_attempt_context(),
|
||||
route_is_loopback=is_loopback_base_url(target.get("base_url")),
|
||||
prompt_tokens_bounded_estimate=bounded_tokens,
|
||||
)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -50,7 +50,7 @@ def normalize_reasoning_effort(value: str, default: str = "medium") -> str:
|
|||
from ouroboros.config import EFFORT_SCALE as _SCALE
|
||||
allowed = set(_SCALE)
|
||||
except Exception:
|
||||
allowed = {"none", "minimal", "low", "medium", "high", "xhigh", "max"}
|
||||
allowed = {"none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra"}
|
||||
v = str(value or "").strip().lower()
|
||||
return v if v in allowed else default
|
||||
|
||||
|
|
|
|||
|
|
@ -32,23 +32,8 @@ class LocalContextTooLargeError(RuntimeError):
|
|||
"""Raised when a local model cannot fit context without silent truncation."""
|
||||
|
||||
|
||||
def _estimate_message_chars(messages: List[Dict[str, Any]]) -> int:
|
||||
from ouroboros.context_budget import IMAGE_BLOCK_CHAR_EQUIVALENT
|
||||
|
||||
total = 0
|
||||
for msg in messages:
|
||||
content = msg.get("content")
|
||||
if isinstance(content, list):
|
||||
for block in content:
|
||||
if not isinstance(block, dict):
|
||||
continue
|
||||
if str(block.get("type") or "") in ("image_url", "image"):
|
||||
total += IMAGE_BLOCK_CHAR_EQUIVALENT
|
||||
continue
|
||||
total += len(str(block.get("text", "")))
|
||||
else:
|
||||
total += len(str(content or ""))
|
||||
return total
|
||||
# Lives beside its proxy constant; the historical private name stays importable.
|
||||
from ouroboros.context_budget import estimate_message_chars as _estimate_message_chars
|
||||
|
||||
|
||||
def _split_markdown_sections(text: str) -> Tuple[str, List[Tuple[str, str]]]:
|
||||
|
|
|
|||
|
|
@ -741,7 +741,13 @@ from ouroboros.loop_acceptance_review import ( # noqa: E402, F401 -- intentiona
|
|||
_record_acceptance_infra_failure,
|
||||
_prior_acceptance_run,
|
||||
_direct_context_fence_state,
|
||||
_disposition_reason_sha256,
|
||||
_refuse_identical_acceptance,
|
||||
_run_task_acceptance_review_once,
|
||||
_total_paid_acceptance_cycles,
|
||||
acceptance_dialogue_history,
|
||||
acceptance_paid_identity,
|
||||
bind_acceptance_paid_identity,
|
||||
)
|
||||
from ouroboros.loop_round_limits import ( # noqa: E402, F401 -- intentional public re-exports
|
||||
_CompactionRoundContext,
|
||||
|
|
@ -834,7 +840,11 @@ from ouroboros.loop_delivery import ( # noqa: E402, F401 -- intentional public
|
|||
_arm_delivery_control,
|
||||
_hold_delivery_for_skill_action,
|
||||
_parse_delivery_control_object,
|
||||
_parse_delivery_control_body,
|
||||
_resolve_delivery_control,
|
||||
_CHILD_ABSORPTION_HOLD_CONTROL,
|
||||
_DELIVERY_HOLD_CONTROLS,
|
||||
_SKILL_ACTION_HOLD_CONTROL,
|
||||
_compose_delivery_suffix,
|
||||
_no_tool_final_answer,
|
||||
)
|
||||
|
|
@ -847,6 +857,7 @@ from ouroboros.loop_forced_finalization import ( # noqa: E402, F401 -- intentio
|
|||
_forced_orphan_note,
|
||||
_claimed_child_dispositions,
|
||||
_undispositioned_children,
|
||||
_undecided_children_listing,
|
||||
_maybe_enforce_child_absorption_gate,
|
||||
_run_forced_children_acceptance,
|
||||
_enforce_swarm_actions,
|
||||
|
|
|
|||
|
|
@ -9,7 +9,8 @@ import logging
|
|||
import pathlib
|
||||
|
||||
from typing import Any, Callable, Dict, List, Optional
|
||||
from ouroboros.outcomes import ACCEPTANCE_ACCEPTED, ACCEPTANCE_BYPASS_REASONS, ACCEPTANCE_BYPASS_REASON_BY_RAIL, ACCEPTANCE_DECISION_STATUSES, ACCEPTANCE_FINALIZED_UNACCEPTED, ACCEPTANCE_REVISION_REQUESTED, REASON_ACCEPTANCE_REVIEW_SKIPPED_DEADLINE_RESERVE, extract_final_answer, turn_has_reviewable_effects
|
||||
from ouroboros.review_cycles import REASON_REVIEW_CYCLES_EXHAUSTED
|
||||
from ouroboros.outcomes import ACCEPTANCE_ACCEPTED, ACCEPTANCE_BYPASS_REASONS, ACCEPTANCE_BYPASS_REASON_BY_RAIL, ACCEPTANCE_DECISION_STATUSES, ACCEPTANCE_FINALIZED_UNACCEPTED, ACCEPTANCE_REVISION_REQUESTED, REASON_ACCEPTANCE_REVIEW_SKIPPED_DEADLINE_RESERVE, REASON_IDENTICAL_ACCEPTANCE_REFUSED, extract_final_answer, turn_has_reviewable_effects
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
from ouroboros.utils import truncate_review_artifact
|
||||
|
||||
|
|
@ -556,6 +557,13 @@ ACCEPTANCE_DECISION_REASONS = (
|
|||
"review_degraded",
|
||||
"fence_reopen_failed",
|
||||
"infra_failure",
|
||||
# The pacing/wallet reason two branches below already STAMP (`pass_reason ==
|
||||
# REASON_REVIEW_CYCLES_EXHAUSTED`); it was missing from the closed set, so a
|
||||
# spent shared cap shipped a reason no reader could validate.
|
||||
REASON_REVIEW_CYCLES_EXHAUSTED,
|
||||
# A-material (2026-08-30): the resubmit carried no changed candidate and no new
|
||||
# obligation disposition, so the recorded verdict was replayed for free.
|
||||
REASON_IDENTICAL_ACCEPTANCE_REFUSED,
|
||||
# Owner Q2A: the forced children_unabsorbed rail runs the panel but cannot
|
||||
# grant a requested improvement pass; the dangling revision terminalizes.
|
||||
"revision_unavailable_on_forced_rail",
|
||||
|
|
|
|||
|
|
@ -144,7 +144,10 @@ def _latest_agent_acceptance_evidence(llm_trace: Dict[str, Any]) -> Dict[str, An
|
|||
|
||||
def _build_host_acceptance_evidence(ctx: _TaskAcceptanceContext) -> Dict[str, Any]:
|
||||
"""Build the one bounded host packet shared by binding and reviewer input."""
|
||||
from ouroboros.review_evidence import build_task_acceptance_evidence
|
||||
from ouroboros.review_evidence import (
|
||||
UNHASHED_ACCEPTANCE_DIALOGUE_HISTORY_KEY,
|
||||
build_task_acceptance_evidence,
|
||||
)
|
||||
|
||||
committed_this_turn = any(
|
||||
isinstance(call, dict)
|
||||
|
|
@ -168,9 +171,28 @@ def _build_host_acceptance_evidence(ctx: _TaskAcceptanceContext) -> Dict[str, An
|
|||
undecided = getattr(ctx.tools._ctx, "_forced_undispositioned_children", None)
|
||||
if isinstance(undecided, list) and undecided:
|
||||
evidence["undispositioned_children"] = undecided
|
||||
# The dialogue so far, so this panel adjudicates knowing what the previous
|
||||
# ones judged instead of re-raising blind. Bounded here rather than by the
|
||||
# packet budget because it is attached AFTER the builder's budget pass — and
|
||||
# it is the one key `task_acceptance_evidence_revision` excludes, so growing
|
||||
# history can never mint a fresh revision (and thus a fresh paid binding).
|
||||
history = acceptance_dialogue_history(ctx.llm_trace)
|
||||
if history:
|
||||
evidence[UNHASHED_ACCEPTANCE_DIALOGUE_HISTORY_KEY] = history
|
||||
return evidence
|
||||
|
||||
|
||||
def _total_paid_acceptance_cycles(ctx: _TaskAcceptanceContext) -> Any:
|
||||
"""Paid acceptance panels this task TREE has already bought, read from the
|
||||
SAME ledger the wallet claim counts (``claimed_cycles``); ``None`` when the
|
||||
projection is unavailable (a descendant that may observe but not initialize)."""
|
||||
from ouroboros.task_results import project_task_acceptance_review_capacity
|
||||
|
||||
return project_task_acceptance_review_capacity(
|
||||
ctx.tools._ctx, task_id=str(ctx.task_id or ""),
|
||||
).get("claimed_cycles")
|
||||
|
||||
|
||||
def _execute_task_acceptance_panel(ctx: _TaskAcceptanceContext) -> Any:
|
||||
"""Perform the one substantive host panel over the pre-bound evidence."""
|
||||
from ouroboros.review_evidence import task_acceptance_evidence_revision
|
||||
|
|
@ -276,8 +298,12 @@ def _execute_task_acceptance_panel(ctx: _TaskAcceptanceContext) -> Any:
|
|||
)
|
||||
duration_sec = round(time.monotonic() - started, 3)
|
||||
try:
|
||||
from ouroboros.review_cycles import review_max_cycles, review_max_cycles_source
|
||||
from ouroboros.utils import append_jsonl, utc_now_iso
|
||||
|
||||
# A panel that just cost money says what bounded it and how many the tree
|
||||
# has bought: "21 paid panels" was invisible until someone summed receipts.
|
||||
_cap = review_max_cycles()
|
||||
append_jsonl(
|
||||
task_pacing.acceptance_timing_events_path(ctx.tools._ctx),
|
||||
{
|
||||
|
|
@ -287,6 +313,9 @@ def _execute_task_acceptance_panel(ctx: _TaskAcceptanceContext) -> Any:
|
|||
"duration_sec": duration_sec,
|
||||
"pass_index": ctx.passes_done,
|
||||
"aggregate_signal": str(result.aggregate_signal or ""),
|
||||
"effective_max_cycles": "unlimited" if _cap is None else _cap,
|
||||
"cycles_source": review_max_cycles_source(),
|
||||
"total_paid_cycles": _total_paid_acceptance_cycles(ctx),
|
||||
},
|
||||
)
|
||||
except Exception:
|
||||
|
|
@ -356,7 +385,7 @@ def _apply_task_acceptance_result(
|
|||
) -> bool:
|
||||
"""Apply one panel result; return whether the agent must take another round."""
|
||||
from ouroboros.review_substrate import (
|
||||
DIALOGUE_CONTINUE,
|
||||
DIALOGUE_TERMINAL_STATUSES,
|
||||
aggregate_dialogue_status,
|
||||
build_improvement_capsule,
|
||||
dissent_findings,
|
||||
|
|
@ -381,14 +410,29 @@ def _apply_task_acceptance_result(
|
|||
rails_line=ctx.rails_line,
|
||||
open_obligations=open_obligations,
|
||||
)
|
||||
# v6.74.0 (A5): the reviewers' typed dialogue judgement, reduced over
|
||||
# ALL contract-valid actors with the panel's own quorum; persisted for
|
||||
# audit on the authoritative run record whatever branch applies below.
|
||||
# v6.74.0 (A5): the reviewers' typed dialogue judgement, reduced over the
|
||||
# CONTRIBUTING actors with the panel's own quorum; persisted for audit on
|
||||
# the authoritative run record whatever branch applies below. `inconclusive`
|
||||
# (no well-formed vote at all) grants the dialogue NO authority: it is not a
|
||||
# terminal verdict and not a licence to continue — the existing non-dialogue
|
||||
# terminals below decide, exactly as they did before the dialogue existed.
|
||||
dialogue = aggregate_dialogue_status(
|
||||
result, quorum=_acceptance_dialogue_quorum(result),
|
||||
)
|
||||
_attach_dialogue_to_host_run(ctx.llm_trace, dialogue)
|
||||
dialogue_terminal = dialogue["status"] != DIALOGUE_CONTINUE
|
||||
dialogue_terminal = dialogue["status"] in DIALOGUE_TERMINAL_STATUSES
|
||||
if reused and getattr(result, "replayed_from_superseded", False):
|
||||
# A run superseded by an evidence revision replays ONLY into the typed
|
||||
# identical-refusal terminal — never into clean-PASS authorization: its
|
||||
# verdict predates the evidence change, so re-accepting would stamp a
|
||||
# stale PASS (and the trace's superseded rows would contradict the
|
||||
# applied decision — the delivery binding could never match). The
|
||||
# refusal is conservative and consistent: nothing new was bought,
|
||||
# nothing stale is re-authorized.
|
||||
return _refuse_identical_acceptance(
|
||||
ctx, result,
|
||||
dialogue=dialogue, dissent=bool(dissent), open_obligations=open_obligations,
|
||||
)
|
||||
if task_acceptance_is_clean(result):
|
||||
ctx.tools._ctx._task_acceptance_reviewed = True
|
||||
_loop()._end_task_acceptance_fence(ctx.tools._ctx, outcome="terminal")
|
||||
|
|
@ -408,6 +452,12 @@ def _apply_task_acceptance_result(
|
|||
ctx.emit_progress("Task acceptance review: PASS (clean acceptance).")
|
||||
return False
|
||||
|
||||
if reused:
|
||||
return _refuse_identical_acceptance(
|
||||
ctx, result,
|
||||
dialogue=dialogue, dissent=bool(dissent), open_obligations=open_obligations,
|
||||
)
|
||||
|
||||
budget_snapshot = task_pacing.build_budget_snapshot(
|
||||
ctx.tools._ctx, profile=ctx.budget_profile,
|
||||
)
|
||||
|
|
@ -421,7 +471,13 @@ def _apply_task_acceptance_result(
|
|||
),
|
||||
ctx=ctx.tools._ctx,
|
||||
)
|
||||
if dialogue_terminal:
|
||||
# A DEGRADED panel (no valid verdict quorum) cannot "judge" the dialogue:
|
||||
# a lone terminal vote from the one contributing slot must NOT shadow the
|
||||
# review_degraded path below, which is the only surface carrying the
|
||||
# per-slot causes and degraded_reasons the v6.70.0 honesty invariant (P1)
|
||||
# requires. Letting the dialogue-terminal branch fire here recorded a false
|
||||
# "reviewer quorum judged" rationale and silently dropped those causes.
|
||||
if dialogue_terminal and str(result.aggregate_signal or "DEGRADED").upper() != "DEGRADED":
|
||||
# v6.74.0 (A5): a reviewer quorum judged the dialogue no longer
|
||||
# actionable (unreachable_here / stable_disagreement). Finalize via
|
||||
# the EXISTING honest path recording BOTH positions in one
|
||||
|
|
@ -654,33 +710,220 @@ def _record_acceptance_infra_failure(ctx: _TaskAcceptanceContext, exc: Exception
|
|||
return False
|
||||
|
||||
|
||||
def _disposition_reason_sha256(reason: Any) -> str:
|
||||
"""Content identity of one obligation-disposition reason; "" when blank.
|
||||
|
||||
Mirrors ``commit_gate.compute_rebuttal_sha256`` on purpose: on BOTH gates an
|
||||
empty rebuttal is not an argument and buys no paid cycle."""
|
||||
import hashlib
|
||||
|
||||
text = str(reason or "").strip()
|
||||
if not text:
|
||||
return ""
|
||||
return hashlib.sha256(text.encode("utf-8", errors="replace")).hexdigest()
|
||||
|
||||
|
||||
def acceptance_paid_identity(candidate_hash: str, llm_trace: Dict[str, Any]) -> str:
|
||||
"""The identity ONE paid acceptance panel is claimed under (A-material).
|
||||
|
||||
``sha256(candidate_hash + the sorted set of nonempty (obligation_id,
|
||||
disposition, sha256(reason)) tuples)``. Exactly two things mint a new paid
|
||||
panel: a changed candidate answer, or an obligation disposition whose content
|
||||
the reviewers have not answered yet. The evidence revision is deliberately NOT
|
||||
in here — every cosmetic tool call moves it, which is how one task bought 21
|
||||
paid panels; it stays what it always was, stale-packet detection for the
|
||||
supersede paths. A disposition with an empty reason contributes nothing.
|
||||
Rows are read live from the agent's own ``acceptance_obligations`` (the
|
||||
``task_acceptance_review`` tool stamps ``status="agent_disposed"`` there)."""
|
||||
import hashlib
|
||||
|
||||
material = sorted({
|
||||
(
|
||||
str(row.get("id") or "").strip(),
|
||||
str(row.get("disposition") or "").strip().lower(),
|
||||
_disposition_reason_sha256(row.get("disposition_reason")),
|
||||
)
|
||||
for row in (llm_trace.get("acceptance_obligations") or [])
|
||||
if isinstance(row, dict)
|
||||
and str(row.get("id") or "").strip()
|
||||
and str(row.get("disposition") or "").strip()
|
||||
and _disposition_reason_sha256(row.get("disposition_reason"))
|
||||
})
|
||||
payload = json.dumps(
|
||||
[str(candidate_hash or ""), [list(item) for item in material]],
|
||||
ensure_ascii=False, separators=(",", ":"),
|
||||
)
|
||||
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def bind_acceptance_paid_identity(
|
||||
review_binding: Dict[str, Any], llm_trace: Dict[str, Any],
|
||||
) -> str:
|
||||
"""Stamp the A-material paid identity onto a freshly built review binding.
|
||||
|
||||
The binding keeps carrying its three hashes (the supersede paths still need
|
||||
the evidence revision); ``paid_identity`` rides ALONGSIDE them and is what the
|
||||
wallet claim and the free-replay lookup key on."""
|
||||
identity = acceptance_paid_identity(
|
||||
str(review_binding.get("candidate_hash") or ""), llm_trace,
|
||||
)
|
||||
review_binding["paid_identity"] = identity
|
||||
return identity
|
||||
|
||||
|
||||
def acceptance_dialogue_history(llm_trace: Dict[str, Any], *, limit: int = 6) -> List[Dict[str, Any]]:
|
||||
"""Bounded per-panel history of the dialogue so far, for the NEXT reviewer.
|
||||
|
||||
Reviewers were adjudicating each round blind to the previous rounds' typed
|
||||
judgement, which is most of why the same finding kept being re-raised. The
|
||||
rows are tiny host facts already recorded on the run records; the caller
|
||||
attaches them to the evidence packet OUTSIDE the hashed material
|
||||
(``review_evidence.UNHASHED_EVIDENCE_KEYS``) so reading the history can never
|
||||
mint a fresh evidence revision — and therefore never a fresh paid binding."""
|
||||
rows: List[Dict[str, Any]] = []
|
||||
for run in (llm_trace.get("review_runs") or []):
|
||||
if not isinstance(run, dict) or str(run.get("authority") or "") != "host_root":
|
||||
continue
|
||||
dialogue = run.get("dialogue") if isinstance(run.get("dialogue"), dict) else {}
|
||||
votes = dialogue.get("votes") if isinstance(dialogue.get("votes"), dict) else {}
|
||||
rows.append({
|
||||
"round": len(rows) + 1,
|
||||
"aggregate_signal": str(run.get("aggregate_signal") or "").upper(),
|
||||
"dialogue_status": str(dialogue.get("status") or ""),
|
||||
"votes": {str(k): len(v or []) for k, v in votes.items()},
|
||||
})
|
||||
obligations = [
|
||||
row for row in (llm_trace.get("acceptance_obligations") or [])
|
||||
if isinstance(row, dict)
|
||||
]
|
||||
if rows:
|
||||
rows[-1]["obligations_new"] = sum(
|
||||
1 for row in obligations if not int(row.get("reopened_count") or 0)
|
||||
)
|
||||
rows[-1]["obligations_re_raised"] = sum(
|
||||
1 for row in obligations if int(row.get("reopened_count") or 0)
|
||||
)
|
||||
return rows[-max(1, int(limit)):]
|
||||
|
||||
|
||||
def _refuse_identical_acceptance(
|
||||
ctx: Any,
|
||||
result: Any,
|
||||
*,
|
||||
dialogue: Dict[str, Any],
|
||||
dissent: bool,
|
||||
open_obligations: List[Dict[str, Any]],
|
||||
) -> bool:
|
||||
"""Terminate a resubmit whose A-material paid identity was already bought.
|
||||
|
||||
The recorded verdict is replayed for FREE and quoted; the improvement capsule
|
||||
is deliberately NOT re-entered. Feeding the note again asks for a round the
|
||||
agent has already answered with nothing new, and every such round shifted the
|
||||
evidence revision into a fresh paid binding — the 21-panel pump. The dialogue
|
||||
record and the decision row still land, so the replay stays auditable."""
|
||||
from ouroboros.outcomes import REASON_IDENTICAL_ACCEPTANCE_REFUSED
|
||||
|
||||
ctx.tools._ctx._task_acceptance_reviewed = True
|
||||
_loop()._end_task_acceptance_fence(ctx.tools._ctx, outcome="terminal")
|
||||
_loop()._mark_root_acceptance_checkpoint(
|
||||
ctx.tools._ctx,
|
||||
ctx.llm_trace,
|
||||
status=str(result.aggregate_signal or "DEGRADED").lower(),
|
||||
pass_index=ctx.passes_done,
|
||||
)
|
||||
_loop()._set_acceptance_decision(ctx.llm_trace, {
|
||||
"status": ACCEPTANCE_FINALIZED_UNACCEPTED,
|
||||
"reason": REASON_IDENTICAL_ACCEPTANCE_REFUSED,
|
||||
"source": "task_acceptance_review",
|
||||
"rationale": (
|
||||
"No new material since the paid panel: neither the candidate answer nor "
|
||||
"any obligation disposition changed. Quoting the recorded verdict "
|
||||
f"({str(result.aggregate_signal or 'DEGRADED').upper()}; dialogue "
|
||||
f"{dialogue['status']}) with {len(open_obligations)} open "
|
||||
"obligation(s); no further round."
|
||||
),
|
||||
"dialogue_status": dialogue["status"],
|
||||
"dialogue_votes": dialogue["votes"],
|
||||
"dissent_noted": dissent,
|
||||
"open_obligations": [str(item.get("id")) for item in open_obligations],
|
||||
})
|
||||
ctx.emit_progress(
|
||||
f"Task acceptance review: {result.aggregate_signal} — identical paid identity "
|
||||
"(no changed answer, no new obligation disposition); the recorded verdict "
|
||||
"stands and no further panel is bought."
|
||||
)
|
||||
return False
|
||||
|
||||
|
||||
def _prior_acceptance_run(
|
||||
tools_ctx: Any, llm_trace: Dict[str, Any], binding_hash: str,
|
||||
tools_ctx: Any,
|
||||
llm_trace: Dict[str, Any],
|
||||
binding_hash: str,
|
||||
*,
|
||||
paid_identity: str = "",
|
||||
) -> Tuple[Dict[str, Any], Optional[Dict[str, Any]]]:
|
||||
"""Locate the authoritative host run already recorded for this binding:
|
||||
"""Locate the authoritative host run already recorded for this submission:
|
||||
first the trace (survives requeue replay), then the process-local
|
||||
``_task_acceptance_seen_bindings`` cache. Returns (cache, prior_run)."""
|
||||
``_task_acceptance_seen_bindings`` cache. Returns (cache, prior_run).
|
||||
|
||||
EITHER identity replays for free: the same binding hash (byte-identical
|
||||
submission, as before) OR the same A-material ``paid_identity`` — unchanged
|
||||
candidate answer and no new obligation disposition — which is the identity the
|
||||
tree's wallet actually bought."""
|
||||
seen_bindings = getattr(tools_ctx, "_task_acceptance_seen_bindings", None)
|
||||
if not isinstance(seen_bindings, dict):
|
||||
seen_bindings = {}
|
||||
tools_ctx._task_acceptance_seen_bindings = seen_bindings
|
||||
identity = str(paid_identity or "")
|
||||
|
||||
def _matches(run: Any) -> bool:
|
||||
return isinstance(run, dict) and (
|
||||
str(run.get("binding_hash") or "") == binding_hash
|
||||
or bool(identity and str(run.get("paid_identity") or "") == identity)
|
||||
)
|
||||
|
||||
prior_run = next(
|
||||
(
|
||||
run for run in reversed(llm_trace.get("review_runs") or [])
|
||||
if isinstance(run, dict)
|
||||
and run.get("authority") == "host_root"
|
||||
and not run.get("superseded_by_revision")
|
||||
and str(run.get("binding_hash") or "") == binding_hash
|
||||
and _matches(run)
|
||||
),
|
||||
None,
|
||||
)
|
||||
cached_run = seen_bindings.get(binding_hash)
|
||||
if (
|
||||
prior_run is None
|
||||
and isinstance(cached_run, dict)
|
||||
and not cached_run.get("superseded_by_revision")
|
||||
):
|
||||
prior_run = cached_run
|
||||
if prior_run is None:
|
||||
prior_run = next(
|
||||
(
|
||||
run for run in reversed(list(seen_bindings.values()))
|
||||
if isinstance(run, dict)
|
||||
and not run.get("superseded_by_revision")
|
||||
and _matches(run)
|
||||
),
|
||||
None,
|
||||
)
|
||||
if prior_run is None and identity:
|
||||
# A run superseded by an evidence revision is stale as a CURRENT
|
||||
# acceptance, but the wallet already bought its A-material. When the
|
||||
# resubmission carries the SAME paid identity (unchanged candidate, no
|
||||
# new nonempty disposition), the recorded verdict replays for free —
|
||||
# otherwise the dispatch claim refuses `binding_dispatch_already_claimed`
|
||||
# and the loop records a synthetic DEGRADED panel instead of the typed
|
||||
# identical-refusal terminal the contract requires. Evidence revision
|
||||
# is stale-DETECTION, never a paid-cycle mint (owner decision 5=A).
|
||||
superseded = next(
|
||||
(
|
||||
run for run in reversed(llm_trace.get("review_runs") or [])
|
||||
if isinstance(run, dict)
|
||||
and run.get("authority") == "host_root"
|
||||
and run.get("superseded_by_revision")
|
||||
and str(run.get("paid_identity") or "") == identity
|
||||
),
|
||||
None,
|
||||
)
|
||||
if superseded is not None:
|
||||
prior_run = dict(superseded)
|
||||
prior_run["replayed_from_superseded"] = True
|
||||
return seen_bindings, prior_run
|
||||
|
||||
|
||||
|
|
@ -876,8 +1119,11 @@ def _run_task_acceptance_review_once(
|
|||
fence_token_or_state=_direct_context_fence_state(tools._ctx, _fence_token),
|
||||
)
|
||||
binding_hash = str(review_ctx.review_binding.get("binding_hash") or "")
|
||||
# A-material: what the tree's wallet actually buys. Stamped onto the
|
||||
# binding before the free-replay lookup and the dispatch claim both read it.
|
||||
paid_identity = bind_acceptance_paid_identity(review_ctx.review_binding, llm_trace)
|
||||
seen_bindings, prior_run = _prior_acceptance_run(
|
||||
tools._ctx, llm_trace, binding_hash,
|
||||
tools._ctx, llm_trace, binding_hash, paid_identity=paid_identity,
|
||||
)
|
||||
reused_result = None
|
||||
if prior_run is not None:
|
||||
|
|
|
|||
|
|
@ -607,8 +607,16 @@ def _prepare_post_tool_budget_context(
|
|||
|
||||
candidate = getattr(tools._ctx, "_delivery_candidate", None)
|
||||
if isinstance(candidate, _loop().DeliveryCandidate):
|
||||
skill_action_pending = (
|
||||
candidate.finalization_control == "skill_action_or_revision_required"
|
||||
hold_control = (
|
||||
candidate.finalization_control
|
||||
if candidate.finalization_control in _loop()._DELIVERY_HOLD_CONTROLS
|
||||
else ""
|
||||
)
|
||||
# The absorption gate stays open while undispositioned children remain:
|
||||
# arming JSON there would recreate the conflicting-instruction round.
|
||||
absorption_gate_open = (
|
||||
hold_control == _loop()._CHILD_ABSORPTION_HOLD_CONTROL
|
||||
and bool(_loop()._undispositioned_children(limit_ctx))
|
||||
)
|
||||
evidence_revision, evidence_fingerprint = _loop()._delivery_evidence_state(
|
||||
tools, limit_ctx, llm_trace,
|
||||
|
|
@ -617,19 +625,26 @@ def _prepare_post_tool_budget_context(
|
|||
candidate.evidence_revision != evidence_revision
|
||||
or candidate.evidence_fingerprint != evidence_fingerprint
|
||||
):
|
||||
_loop()._arm_delivery_control(
|
||||
tools,
|
||||
limit_ctx,
|
||||
llm_trace,
|
||||
control="effect_revision_required",
|
||||
)
|
||||
elif skill_action_pending:
|
||||
if absorption_gate_open:
|
||||
_loop()._hold_delivery_for_skill_action(
|
||||
tools, llm_trace, control=_loop()._CHILD_ABSORPTION_HOLD_CONTROL,
|
||||
)
|
||||
else:
|
||||
_loop()._arm_delivery_control(
|
||||
tools,
|
||||
limit_ctx,
|
||||
llm_trace,
|
||||
control="effect_revision_required",
|
||||
)
|
||||
elif hold_control == _loop()._SKILL_ACTION_HOLD_CONTROL:
|
||||
_loop()._arm_delivery_control(
|
||||
tools,
|
||||
limit_ctx,
|
||||
llm_trace,
|
||||
control="skill_revision_required",
|
||||
)
|
||||
# An absorption hold with unchanged evidence keeps holding: only
|
||||
# dispositions close the gate, and a disposition changes evidence.
|
||||
# Cross-model fallback can adopt a different route during this round.
|
||||
limit_ctx.active_model = active_model
|
||||
limit_ctx.active_use_local = active_use_local
|
||||
|
|
|
|||
|
|
@ -58,6 +58,18 @@ class DeliveryCandidate:
|
|||
model_text: str = ""
|
||||
|
||||
|
||||
# Action-gate holds: a gate closable ONLY by a tool call (skill lifecycle
|
||||
# action, child-result disposition) retains the candidate WITHOUT arming the
|
||||
# JSON-only control instruction — the instruction would conflict with the
|
||||
# required tool call. Gates closable by a reconsidered answer arm normally.
|
||||
_SKILL_ACTION_HOLD_CONTROL = "skill_action_or_revision_required"
|
||||
_CHILD_ABSORPTION_HOLD_CONTROL = "child_absorption_or_revision_required"
|
||||
_DELIVERY_HOLD_CONTROLS = frozenset({
|
||||
_SKILL_ACTION_HOLD_CONTROL,
|
||||
_CHILD_ABSORPTION_HOLD_CONTROL,
|
||||
})
|
||||
|
||||
|
||||
def _swarm_handoff_attempt(ctx: Any) -> Dict[str, Any]:
|
||||
attempt = getattr(ctx, "_swarm_handoff_attempt", None)
|
||||
return dict(attempt) if isinstance(attempt, dict) else {}
|
||||
|
|
@ -511,13 +523,19 @@ def _arm_delivery_control(
|
|||
def _hold_delivery_for_skill_action(
|
||||
tools: ToolRegistry,
|
||||
llm_trace: Dict[str, Any],
|
||||
*,
|
||||
control: str = _SKILL_ACTION_HOLD_CONTROL,
|
||||
) -> None:
|
||||
"""Retain the answer while an unresolved skill lifecycle gate requires action."""
|
||||
"""Retain the answer while an unresolved action gate requires a tool call.
|
||||
|
||||
``control`` names the open gate; it must stay within
|
||||
``_DELIVERY_HOLD_CONTROLS`` so both hold readers recognize the state.
|
||||
"""
|
||||
|
||||
candidate = getattr(tools._ctx, "_delivery_candidate", None)
|
||||
if not isinstance(candidate, _loop().DeliveryCandidate):
|
||||
return
|
||||
candidate.finalization_control = "skill_action_or_revision_required"
|
||||
candidate.finalization_control = control
|
||||
candidate.repair_attempted = False
|
||||
tools._ctx._delivery_control_required = False
|
||||
_loop()._publish_delivery_candidate(tools, candidate, llm_trace)
|
||||
|
|
@ -547,13 +565,51 @@ def _parse_delivery_control_object(
|
|||
|
||||
try:
|
||||
payload = json.loads(raw, object_pairs_hook=_unique_object)
|
||||
except (TypeError, ValueError, json.JSONDecodeError):
|
||||
except (TypeError, ValueError, json.JSONDecodeError, RecursionError):
|
||||
# RecursionError: a degenerate deeply-nested blob (repetition-loop
|
||||
# model output) must classify as not-a-control, not crash the round.
|
||||
return None, duplicate_protocol_key
|
||||
if not isinstance(payload, dict):
|
||||
return None, False
|
||||
return payload, False
|
||||
|
||||
|
||||
def _parse_delivery_control_body(
|
||||
raw: str,
|
||||
) -> Tuple[Optional[Dict[str, Any]], bool, bool]:
|
||||
"""Normalize a response body and locate its delivery-control object.
|
||||
|
||||
Returns ``(parsed, duplicate_protocol_key, embedded)``. Normalization
|
||||
strips one whole-body markdown fence (shared with
|
||||
``observability._is_delivery_control_payload``). ``embedded`` is True only
|
||||
when the protocol object sits as a balanced trailing JSON object carrying
|
||||
the ``delivery_control`` key at the very END of surrounding prose — a
|
||||
protocol attempt mixed with text, never a valid control. A control object
|
||||
quoted MID-prose is NOT matched and stays prose (disclosed residual)."""
|
||||
from ouroboros.observability import strip_protocol_fence
|
||||
|
||||
body = strip_protocol_fence(raw)
|
||||
parsed, duplicate_protocol_key = _parse_delivery_control_object(body)
|
||||
if duplicate_protocol_key or isinstance(parsed, dict):
|
||||
return parsed, duplicate_protocol_key, False
|
||||
# Trailing scan: ONE O(n) string-aware pass over the body (fenced and
|
||||
# double-fenced tails peeled, duplicate keys flagged, RecursionError
|
||||
# degraded, bounded line-anchor retries after an unbalanced prose brace
|
||||
# or quote). The extractor is key-agnostic; the protocol judgment stays
|
||||
# HERE: only a trailing object carrying `delivery_control` at its top
|
||||
# level (or a duplicated protocol key) is an embedded protocol attempt.
|
||||
from ouroboros.utils import extract_trailing_json_object
|
||||
|
||||
_prefix, tail_parsed, tail_duplicate = extract_trailing_json_object(
|
||||
body, duplicate_flag_keys=("delivery_control", "full_answer"),
|
||||
)
|
||||
if tail_duplicate:
|
||||
return None, True, True
|
||||
if isinstance(tail_parsed, dict) and "delivery_control" in tail_parsed:
|
||||
return tail_parsed, False, True
|
||||
return None, False, False
|
||||
|
||||
|
||||
def _resolve_delivery_control(
|
||||
content: Any,
|
||||
tools: ToolRegistry,
|
||||
|
|
@ -567,10 +623,10 @@ def _resolve_delivery_control(
|
|||
if not isinstance(candidate, _loop().DeliveryCandidate):
|
||||
return "fresh", _loop()._extract_plain_text_from_content(content)
|
||||
raw = _loop()._extract_plain_text_from_content(content).strip()
|
||||
parsed, duplicate_protocol_key = _loop()._parse_delivery_control_object(raw)
|
||||
# ANY parsed object carrying the protocol key is control intent,
|
||||
# regardless of verb/value — an unknown verb is a mangled protocol
|
||||
# attempt, never prose (raw JSON leaked to chat); validity judged below.
|
||||
parsed, duplicate_protocol_key, embedded_protocol = _parse_delivery_control_body(raw)
|
||||
# ANY parsed object carrying the protocol key is control intent, whatever
|
||||
# the verb or placement — a mangled protocol attempt is never prose (raw
|
||||
# JSON leaked to chat); validity judged below.
|
||||
is_control_intent = duplicate_protocol_key or (
|
||||
isinstance(parsed, dict) and "delivery_control" in parsed
|
||||
)
|
||||
|
|
@ -581,12 +637,12 @@ def _resolve_delivery_control(
|
|||
# required latch. The candidate's typed control state is authoritative.
|
||||
required = True
|
||||
tools._ctx._delivery_control_required = True
|
||||
elif candidate.finalization_control == "skill_action_or_revision_required":
|
||||
# Preserve the historical bounded skill gate: an actual tool
|
||||
# action or a reconsidered full prose answer may proceed, but a
|
||||
# typed keep cannot acknowledge the gate. No delivery JSON prompt
|
||||
# before the action — it would conflict with the instruction to
|
||||
# call the skill lifecycle tool.
|
||||
elif candidate.finalization_control in _DELIVERY_HOLD_CONTROLS:
|
||||
# Bounded action gates (skill lifecycle, child absorption): a tool
|
||||
# action or a reconsidered full prose answer may proceed; a typed
|
||||
# keep cannot acknowledge the gate and no JSON prompt rides the
|
||||
# action round. A typed control attempt escalates to the ONE
|
||||
# replace-required literal for BOTH holds (plan-rejected widening).
|
||||
if not is_control_intent:
|
||||
return "fresh", _loop()._extract_plain_text_from_content(content)
|
||||
candidate.finalization_control = "skill_revision_required"
|
||||
|
|
@ -608,7 +664,12 @@ def _resolve_delivery_control(
|
|||
selected = str(parsed.get("delivery_control") or "") if isinstance(parsed, dict) else ""
|
||||
valid = False
|
||||
replacement = ""
|
||||
if selected == "keep" and set(parsed) == {"delivery_control"}:
|
||||
if embedded_protocol:
|
||||
# A trailing prose-embedded object is a protocol ATTEMPT, never a
|
||||
# valid control: honoring it would leak the raw object or drop the
|
||||
# prose half (the default error states the exact-object rule).
|
||||
pass
|
||||
elif selected == "keep" and set(parsed) == {"delivery_control"}:
|
||||
valid = _delivery_keep_allowed(
|
||||
candidate, evidence_revision, evidence_fingerprint,
|
||||
)
|
||||
|
|
@ -731,7 +792,11 @@ def _no_tool_final_answer(
|
|||
tools, limit_ctx, content, messages, emit_progress, llm_trace,
|
||||
)
|
||||
if absorption_result == "continue":
|
||||
_loop()._arm_delivery_control(tools, limit_ctx, llm_trace)
|
||||
# Child absorption is closable only by disposition tool calls: hold —
|
||||
# arming the JSON-only instruction would contradict the reminder.
|
||||
_hold_delivery_for_skill_action(
|
||||
tools, llm_trace, control=_CHILD_ABSORPTION_HOLD_CONTROL,
|
||||
)
|
||||
return None
|
||||
if absorption_result is not None:
|
||||
return absorption_result
|
||||
|
|
|
|||
|
|
@ -316,6 +316,19 @@ def _undispositioned_children(ctx: _RoundLimitContext) -> list[Dict[str, Any]]:
|
|||
return []
|
||||
|
||||
|
||||
def _undecided_children_listing(undecided: list[Dict[str, Any]]) -> str:
|
||||
"""Bounded ``id [status] sha256`` listing shared by the absorption
|
||||
reminder and the forced-finalization prompt."""
|
||||
|
||||
from ouroboros.tools.join_ledger import _child_result_sha256
|
||||
|
||||
return "; ".join(
|
||||
f"{c.get('task_id') or c.get('id') or '?'} [{c.get('status') or 'unknown'}] "
|
||||
f"sha256={_child_result_sha256(c)}"
|
||||
for c in undecided[:10]
|
||||
)
|
||||
|
||||
|
||||
def _maybe_enforce_child_absorption_gate(
|
||||
tools: ToolRegistry,
|
||||
limit_ctx: _RoundLimitContext,
|
||||
|
|
@ -331,13 +344,7 @@ def _maybe_enforce_child_absorption_gate(
|
|||
tools._ctx._child_absorption_reminded = True
|
||||
if content and str(content).strip():
|
||||
messages.append({"role": "assistant", "content": content})
|
||||
from ouroboros.tools.join_ledger import _child_result_sha256
|
||||
|
||||
listed = "; ".join(
|
||||
f"{c.get('task_id') or c.get('id') or '?'} [{c.get('status') or 'unknown'}] "
|
||||
f"sha256={_child_result_sha256(c)}"
|
||||
for c in undecided[:10]
|
||||
)
|
||||
listed = _undecided_children_listing(undecided)
|
||||
reminder = (
|
||||
"[CHILD_ABSORPTION_REQUIRED]\n"
|
||||
"You have child result(s) without a current exact-hash disposition: "
|
||||
|
|
@ -354,20 +361,24 @@ def _maybe_enforce_child_absorption_gate(
|
|||
emit_progress("Child absorption reminder injected before final response.")
|
||||
llm_trace["reasoning_notes"].append("Child absorption reminder injected before final response.")
|
||||
return "continue"
|
||||
# Fresh snapshot for the forced prompt: child statuses may have flipped
|
||||
# since the reminder round; the model must state CURRENT statuses.
|
||||
undecided = _undispositioned_children(limit_ctx)
|
||||
text, usage, forced_trace = _loop()._forced_final_answer(
|
||||
limit_ctx,
|
||||
prompt=(
|
||||
"[FINALIZE_WITH_UNABSORBED_CHILDREN]\n"
|
||||
"You still have child results without exact dispositions and already received one "
|
||||
"child-absorption reminder. Produce an honest best-effort final answer now; name the "
|
||||
"unabsorbed or unfinished children explicitly."
|
||||
"unabsorbed or unfinished children explicitly. Current child state: "
|
||||
f"{_undecided_children_listing(undecided)}."
|
||||
),
|
||||
fallback_text="⚠️ Finalized best-effort with undispositioned child results.",
|
||||
reason_code="children_unabsorbed",
|
||||
)
|
||||
_loop()._merge_finalization_trace(llm_trace, forced_trace)
|
||||
_run_forced_children_acceptance(
|
||||
tools, limit_ctx, undecided, text, messages, emit_progress, llm_trace,
|
||||
tools, limit_ctx, text, messages, emit_progress, llm_trace,
|
||||
)
|
||||
return text, usage, llm_trace
|
||||
|
||||
|
|
@ -375,7 +386,6 @@ def _maybe_enforce_child_absorption_gate(
|
|||
def _run_forced_children_acceptance(
|
||||
tools: ToolRegistry,
|
||||
limit_ctx: _RoundLimitContext,
|
||||
undecided: list[Dict[str, Any]],
|
||||
text: str,
|
||||
messages: List[Dict[str, Any]],
|
||||
emit_progress: Callable[[str], None],
|
||||
|
|
@ -397,6 +407,9 @@ def _run_forced_children_acceptance(
|
|||
try:
|
||||
from ouroboros.tools.join_ledger import _child_result_sha256
|
||||
|
||||
# Fresh debt adjacent to the panel's own fresh subtree read: a child
|
||||
# may settle across the forced call — one packet, one moment.
|
||||
undecided = _undispositioned_children(limit_ctx)
|
||||
debt = [
|
||||
{
|
||||
"task_id": str(c.get("task_id") or c.get("id") or ""),
|
||||
|
|
@ -860,27 +873,31 @@ def _resolve_forced_delivery_control(
|
|||
if not armed:
|
||||
return extracted, ""
|
||||
tools_ctx._delivery_control_required = False
|
||||
parsed, duplicate_protocol_key = _loop()._parse_delivery_control_object(extracted)
|
||||
from ouroboros.observability import strip_protocol_fence
|
||||
|
||||
parsed, duplicate_protocol_key, embedded_protocol = _loop()._parse_delivery_control_body(extracted)
|
||||
# Protocol intent: any parsed object with the protocol key (unknown verb =
|
||||
# broken control, never prose), or JSON-looking text that fails to parse (a
|
||||
# mangled protocol attempt under the armed latch — the candidate is the answer).
|
||||
# broken control; a trailing prose-embedded object counts), or JSON-looking
|
||||
# text after the shared fence-strip that fails to parse under the latch.
|
||||
protocol_intent = duplicate_protocol_key or (
|
||||
("delivery_control" in parsed)
|
||||
if isinstance(parsed, dict)
|
||||
else extracted.lstrip().startswith("{")
|
||||
else strip_protocol_fence(extracted).startswith("{")
|
||||
)
|
||||
if not protocol_intent:
|
||||
# An ordinary prose answer under an armed latch: the fresh text stands.
|
||||
# Ordinary prose under an armed latch stands (a control object quoted
|
||||
# MID-prose is the disclosed residual).
|
||||
return extracted, ""
|
||||
selected = str(parsed.get("delivery_control") or "") if isinstance(parsed, dict) else ""
|
||||
if selected == "replace" and set(parsed) == {"delivery_control", "full_answer"}:
|
||||
replacement = parsed.get("full_answer")
|
||||
if isinstance(replacement, str) and replacement.strip():
|
||||
return replacement, ""
|
||||
elif selected == "keep" and set(parsed) == {"delivery_control"} and candidate is not None:
|
||||
return candidate.full_text, ""
|
||||
# Malformed/duplicate/invalid control: preserve the retained candidate (or,
|
||||
# with none retained, let the caller's fallback text stand) and say so.
|
||||
if not embedded_protocol:
|
||||
selected = str(parsed.get("delivery_control") or "") if isinstance(parsed, dict) else ""
|
||||
if selected == "replace" and set(parsed) == {"delivery_control", "full_answer"}:
|
||||
replacement = parsed.get("full_answer")
|
||||
if isinstance(replacement, str) and replacement.strip():
|
||||
return replacement, ""
|
||||
elif selected == "keep" and set(parsed) == {"delivery_control"} and candidate is not None:
|
||||
return candidate.full_text, ""
|
||||
# Malformed/duplicate/prose-embedded/invalid control: preserve the retained
|
||||
# candidate (with none retained, the caller's fallback stands) and say so.
|
||||
return (
|
||||
candidate.full_text if candidate is not None else "",
|
||||
REASON_DELIVERY_CONTROL_DEGRADED,
|
||||
|
|
|
|||
|
|
@ -24,7 +24,7 @@ from ouroboros.outcomes import REASON_OWNER_REQUESTED_FINALIZATION
|
|||
from ouroboros.pricing import estimate_cost_optional
|
||||
from ouroboros.task_finalization import TERMINAL_ORIGIN_HOST_SALVAGE
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
from supervisor.owner_stop import _narrow_round_deadline, _owner_stop_control_is_current, _owner_stop_window_elapsed
|
||||
from supervisor.owner_stop import _narrow_round_deadline, _owner_stop_control_is_current, _owner_stop_window_elapsed, handle_finalize_now_entry # noqa: F401 -- _owner_stop_control_is_current stays a facade surface
|
||||
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
|
|
@ -92,7 +92,7 @@ def _drain_incoming_messages(
|
|||
break
|
||||
|
||||
if drive_root is not None and task_id:
|
||||
from ouroboros.owner_mailbox import KIND_FINALIZE_NOW, KIND_HURRY, KIND_OWNER_TEXT, KIND_TASK_MESSAGE, acknowledge_transcript_entry, deliver_task_message, drain_owner_entries
|
||||
from ouroboros.owner_mailbox import KIND_FINALIZE_NOW, KIND_HURRY, KIND_OWNER_TEXT, KIND_QUIZ_ANSWER, KIND_TASK_MESSAGE, acknowledge_transcript_entry, deliver_quiz_answer, deliver_task_message, drain_owner_entries
|
||||
|
||||
if owner_ctx:
|
||||
owner_ctx._loop_mailbox_seen_ids = _owner_msg_seen
|
||||
|
|
@ -100,36 +100,7 @@ def _drain_incoming_messages(
|
|||
for entry in drain_owner_entries(drive_root, task_id, _owner_msg_seen, attempt):
|
||||
kind = entry.get("kind") or KIND_OWNER_TEXT
|
||||
if kind == KIND_FINALIZE_NOW:
|
||||
text = str(entry.get("text") or "deadline")
|
||||
first_line = text.splitlines()[0].strip() if text else ""
|
||||
if first_line == REASON_OWNER_REQUESTED_FINALIZATION:
|
||||
if not _owner_stop_control_is_current(
|
||||
owner_ctx,
|
||||
drive_root,
|
||||
task_id,
|
||||
str(entry.get("msg_id") or ""),
|
||||
):
|
||||
continue
|
||||
# Owner-stop budget starts at delivery; first drain wins.
|
||||
if not _loop()._mark_owner_stop_control_drained(
|
||||
owner_ctx, drive_root, task_id,
|
||||
):
|
||||
continue
|
||||
if not _owner_stop_control_is_current(
|
||||
owner_ctx,
|
||||
drive_root,
|
||||
task_id,
|
||||
str(entry.get("msg_id") or ""),
|
||||
):
|
||||
continue
|
||||
else:
|
||||
opened = parse_deadline_ts(entry.get("ts"))
|
||||
if opened is not None:
|
||||
controls["finalize_deadline_ts"] = (
|
||||
opened.timestamp()
|
||||
+ task_pacing.effective_finalization_reserve_sec(owner_ctx)
|
||||
)
|
||||
controls["finalize_now"] = text
|
||||
handle_finalize_now_entry(entry, owner_ctx, drive_root, task_id, controls)
|
||||
continue
|
||||
if kind == KIND_HURRY:
|
||||
# HQ1 no-chat contract (§19.7.2 item 6): a typed hurry
|
||||
|
|
@ -146,6 +117,10 @@ def _drain_incoming_messages(
|
|||
deliver_task_message(entry, task_id, event_queue, lambda text: _loop()._append_or_merge_user_message(messages, text))
|
||||
acknowledge_transcript_entry(drive_root, task_id, entry)
|
||||
continue
|
||||
if kind == KIND_QUIZ_ANSWER:
|
||||
deliver_quiz_answer(entry, task_id, event_queue, lambda text: _loop()._append_or_merge_user_message(messages, text))
|
||||
acknowledge_transcript_entry(drive_root, task_id, entry)
|
||||
continue
|
||||
_loop()._record_owner_directive(
|
||||
owner_ctx,
|
||||
source="owner_mailbox",
|
||||
|
|
|
|||
|
|
@ -2,17 +2,16 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import concurrent.futures
|
||||
import contextvars
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import pathlib
|
||||
import time
|
||||
import concurrent.futures
|
||||
import contextvars
|
||||
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||
from typing import Any, Callable, Dict, List, Optional
|
||||
|
||||
import logging
|
||||
|
||||
from ouroboros.config import (
|
||||
NESTED_SETTLEMENT_MARGIN_SEC,
|
||||
get_finalization_grace_sec,
|
||||
|
|
@ -21,13 +20,19 @@ from ouroboros.config import (
|
|||
from ouroboros.deadline_utils import deadline_remaining_sec
|
||||
from ouroboros.observability import new_call_id, persist_call
|
||||
from ouroboros.tool_capabilities import (
|
||||
READ_ONLY_PARALLEL_TOOLS,
|
||||
PARALLEL_SAFE_ENQUEUE_TOOLS,
|
||||
FOREGROUND_MUTATIVE_TOOLS,
|
||||
PARALLEL_SAFE_ENQUEUE_TOOLS,
|
||||
READ_ONLY_PARALLEL_TOOLS,
|
||||
REVIEWED_MUTATIVE_TOOLS,
|
||||
STATEFUL_BROWSER_TOOLS,
|
||||
UNTRUNCATED_TOOL_RESULTS as _UNTRUNCATED_TOOL_RESULTS,
|
||||
)
|
||||
from ouroboros.tool_capabilities import (
|
||||
UNTRUNCATED_REPO_READ_PATHS as _UNTRUNCATED_REPO_READ_PATHS,
|
||||
)
|
||||
from ouroboros.tool_capabilities import (
|
||||
UNTRUNCATED_TOOL_RESULTS as _UNTRUNCATED_TOOL_RESULTS,
|
||||
)
|
||||
from ouroboros.tool_capabilities import (
|
||||
tool_result_limit as _tool_result_limit,
|
||||
)
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
|
|
@ -101,12 +106,18 @@ def _attach_late_tool_settlement(
|
|||
task_id: str,
|
||||
tool_call_id: str,
|
||||
correlation: Dict[str, Any],
|
||||
on_settled: Optional[Callable[[], None]] = None,
|
||||
) -> None:
|
||||
"""Close the cognitive lease when a timed-out worker finally settles."""
|
||||
tool_ctx = getattr(tools, "_ctx", None)
|
||||
event_queue = getattr(tool_ctx, "event_queue", None)
|
||||
|
||||
def _settled(_future: Any) -> None:
|
||||
if on_settled is not None:
|
||||
try:
|
||||
on_settled()
|
||||
except Exception:
|
||||
log.debug("Late tool cleanup failed", exc_info=True)
|
||||
emit_cognitive_operation_event(
|
||||
event_queue,
|
||||
task_id=task_id,
|
||||
|
|
@ -177,6 +188,13 @@ def _with_correlation(payload: Dict[str, Any], correlation: Dict[str, Any], *, t
|
|||
|
||||
|
||||
_PER_CALL_TIMEOUT_TOOLS = ("run_command", "run_script")
|
||||
# Process tools whose handler publishes TYPED process facts (exit_code, POSIX
|
||||
# signal name, duration_ms — plus resolved_runtime when the interpreter
|
||||
# resolver substituted the executable) through the thread-local channel in
|
||||
# ouroboros.tools.process_facts. Typed facts take PRECEDENCE over the regex harvest
|
||||
# below (_EXIT_CODE_RE/_SIGNAL_RE), which remains as the read-fallback for
|
||||
# records that lack typed meta (older traces, prose-only paths).
|
||||
_PROCESS_META_TOOLS = frozenset({"run_command", "run_script"})
|
||||
# Structural ordering margin: the outer cap sits this far above the requested
|
||||
# per-call timeout so the handler's own (cleanly-messaged) subprocess timeout
|
||||
# fires first, before the outer thread-kill. Not a wait duration — a race margin.
|
||||
|
|
@ -501,6 +519,40 @@ def _tool_result_fields(result: ToolResult) -> Dict[str, Any]:
|
|||
"tool_result_meta": dict(result.meta),
|
||||
}
|
||||
|
||||
def _execute_browser_tool_bound(
|
||||
tools: ToolRegistry,
|
||||
tc: Dict[str, Any],
|
||||
drive_logs: pathlib.Path,
|
||||
task_id: str,
|
||||
generation: Any,
|
||||
) -> Dict[str, Any]:
|
||||
"""The stateful (browser) submit wrapper: the call is BOUND to the browser
|
||||
generation the main thread saw at submit time. A call that only STARTS
|
||||
after its generation was retired (its own timeout fired before the worker
|
||||
reached the tool body) must not touch the replacement — its result goes to
|
||||
the settlement void anyway, so it refuses instead of executing."""
|
||||
tool_ctx = getattr(tools, "_ctx", None)
|
||||
current = getattr(tool_ctx, "browser_state", None) if tool_ctx is not None else None
|
||||
if generation is not None and current is not None and current is not generation:
|
||||
return {
|
||||
"tool_call_id": tc.get("id"),
|
||||
"fn_name": str(tc.get("function", {}).get("name") or ""),
|
||||
"result": ("⚠️ BROWSER_SESSION_RETIRED: this call timed out before it "
|
||||
"started and its browser generation was retired; nothing ran."),
|
||||
"is_error": True,
|
||||
"args_for_log": {},
|
||||
"is_code_tool": False,
|
||||
}
|
||||
if tool_ctx is not None and generation is not None:
|
||||
# Pin the expectation for the tool body: _ensure_browser re-checks it
|
||||
# at the exact moment it captures the generation, which closes the
|
||||
# check-then-run window (safety checks between here and the handler
|
||||
# can be long). _ensure_browser's own legitimate replacements update
|
||||
# the pin, so a mid-call error after such a replacement is never
|
||||
# mistaken for a timeout retirement.
|
||||
setattr(tool_ctx, "_active_browser_generation", generation)
|
||||
return _execute_single_tool(tools, tc, drive_logs, task_id)
|
||||
|
||||
|
||||
def _execute_single_tool(
|
||||
tools: ToolRegistry,
|
||||
|
|
@ -571,6 +623,15 @@ def _execute_single_tool(
|
|||
|
||||
args_for_log = sanitize_tool_args_for_log(fn_name, args if isinstance(args, dict) else {})
|
||||
|
||||
typed_process_meta = fn_name in _PROCESS_META_TOOLS
|
||||
if typed_process_meta:
|
||||
# Defensive: drop any stale thread-local facts before dispatch so a
|
||||
# process-tool path that runs NO process (arg errors, blocks) can never
|
||||
# inherit a previous call's measurements.
|
||||
from ouroboros.tools.process_facts import consume_last_process_facts
|
||||
|
||||
consume_last_process_facts()
|
||||
|
||||
tool_ok = True
|
||||
try:
|
||||
tool_result = tools.execute_result(fn_name, args)
|
||||
|
|
@ -596,6 +657,22 @@ def _execute_single_tool(
|
|||
**_typed_result_metadata(fn_name, result, is_error, tool_result),
|
||||
**_tool_result_fields(tool_result),
|
||||
}
|
||||
if typed_process_meta:
|
||||
# R5 (node-runtime sprint): merge the handler's TYPED process facts into
|
||||
# the call's result_meta. When a typed publication exists it owns the
|
||||
# WHOLE fact family — including the ABSENCE of a member (a typed
|
||||
# publication without ``signal`` means the child was not signal-killed).
|
||||
# It carries the members the ToolResult meta does not (duration_ms,
|
||||
# resolved_runtime) and takes precedence over the producer-meta copy of
|
||||
# the shared ones; records with no typed publication keep the
|
||||
# ToolResult-meta facts from _typed_result_metadata unchanged.
|
||||
from ouroboros.tools.process_facts import PROCESS_FACT_KEYS
|
||||
|
||||
process_facts = consume_last_process_facts()
|
||||
if process_facts:
|
||||
for stale_key in PROCESS_FACT_KEYS:
|
||||
result_meta.pop(stale_key, None)
|
||||
result_meta.update(process_facts)
|
||||
|
||||
trace_ref = {}
|
||||
try:
|
||||
|
|
@ -676,6 +753,16 @@ class StatefulToolExecutor:
|
|||
self._executor.shutdown(wait=False, cancel_futures=True)
|
||||
self._executor = None
|
||||
|
||||
def retire(self):
|
||||
"""Detach from the sticky thread WITHOUT cancelling queued work.
|
||||
|
||||
The timeout path queues the browser-generation cleanup BEHIND the
|
||||
hung call on this same thread; reset()'s cancel_futures would cancel
|
||||
exactly that cleanup and leak the browser."""
|
||||
if self._executor is not None:
|
||||
self._executor.shutdown(wait=False, cancel_futures=False)
|
||||
self._executor = None
|
||||
|
||||
def shutdown(self, wait=True, cancel_futures=False):
|
||||
"""Shutdown the sticky executor."""
|
||||
if self._executor is not None:
|
||||
|
|
@ -796,6 +883,7 @@ def _execute_with_timeout(
|
|||
use_stateful = stateful_executor and fn_name in STATEFUL_BROWSER_TOOLS
|
||||
started_at = time.perf_counter()
|
||||
correlation = _tool_correlation(tools)
|
||||
tool_ctx = getattr(tools, "_ctx", None)
|
||||
args_for_log = {}
|
||||
try:
|
||||
args = json.loads(tc["function"]["arguments"] or "{}")
|
||||
|
|
@ -813,7 +901,16 @@ def _execute_with_timeout(
|
|||
}, correlation, tool_call_id=tool_call_id))
|
||||
|
||||
if use_stateful:
|
||||
future = stateful_executor.submit(_execute_single_tool, tools, tc, drive_logs, task_id)
|
||||
assert stateful_executor is not None
|
||||
# The generation this call is bound to, captured by the SUBMITTING
|
||||
# thread: if the call's own timeout retires it before the worker even
|
||||
# reaches the tool body, the wrapper refuses instead of letting the
|
||||
# abandoned call build a session in the NEXT command's state.
|
||||
submit_generation = getattr(tool_ctx, "browser_state", None)
|
||||
future = stateful_executor.submit(
|
||||
_execute_browser_tool_bound, tools, tc, drive_logs, task_id,
|
||||
submit_generation,
|
||||
)
|
||||
try:
|
||||
result = future.result(timeout=timeout_sec)
|
||||
result_meta = result.get("result_meta") or {}
|
||||
|
|
@ -833,11 +930,46 @@ def _execute_with_timeout(
|
|||
}, correlation, tool_call_id=tool_call_id))
|
||||
return result
|
||||
except (TimeoutError, concurrent.futures.TimeoutError):
|
||||
settlement_target = future
|
||||
on_settled: Optional[Callable[[], None]] = None
|
||||
if tool_ctx is not None: # stateful branch: use_stateful is already true here
|
||||
from ouroboros.tools.browser import (
|
||||
_detach_browser,
|
||||
cleanup_browser_handles,
|
||||
)
|
||||
|
||||
# CROSS-THREAD INVARIANT (#409): Playwright objects are bound
|
||||
# to the worker thread that created them. This main-thread
|
||||
# path may only DETACH (retire) the generation — never call
|
||||
# close()/stop() here (a cross-thread stop() kills the driver
|
||||
# under the hung worker and turns its dispatcher wait into a
|
||||
# CPU busy-loop). The close is QUEUED on the retiring
|
||||
# executor behind the hung call, so it runs on the owning
|
||||
# worker thread whenever that call settles — including the
|
||||
# already-settled race, where the queued task runs at once on
|
||||
# that same thread (a done-callback would have run HERE, on
|
||||
# the main thread, exactly the cross-thread close this path
|
||||
# exists to prevent).
|
||||
retired_generation, _fresh = _detach_browser(tool_ctx)
|
||||
try:
|
||||
settlement_target = stateful_executor.submit(
|
||||
cleanup_browser_handles, retired_generation,
|
||||
)
|
||||
except Exception:
|
||||
# Degenerate fallback (executor already dead): settle on
|
||||
# the tool future and close best-effort in its callback.
|
||||
log.debug("cleanup submit failed; falling back to done-callback",
|
||||
exc_info=True)
|
||||
settlement_target = future
|
||||
on_settled = lambda: cleanup_browser_handles(retired_generation) # noqa: E731
|
||||
_attach_late_tool_settlement(
|
||||
tools, future, task_id=task_id, tool_call_id=tool_call_id,
|
||||
tools, settlement_target, task_id=task_id, tool_call_id=tool_call_id,
|
||||
correlation={**correlation, "tool": fn_name},
|
||||
on_settled=on_settled,
|
||||
)
|
||||
stateful_executor.reset()
|
||||
# retire(), not reset(): cancel_futures would cancel exactly the
|
||||
# queued cleanup task this path just submitted.
|
||||
stateful_executor.retire()
|
||||
reset_msg = "Browser state has been reset. "
|
||||
timeout_result = _make_timeout_result(
|
||||
fn_name, tool_call_id, is_code_tool, tc, drive_logs,
|
||||
|
|
|
|||
|
|
@ -84,6 +84,21 @@ def _installer_env(env_root: pathlib.Path, *, ecosystem: str = "") -> Dict[str,
|
|||
"npm_config_cache": str(cache_dir / "npm"), "npm_config_userconfig": str(env_root / "npmrc"),
|
||||
"CARGO_HOME": str(env_root / "cargo" / "home"), "CARGO_TARGET_DIR": str(env_root / "cargo" / "target"),
|
||||
})
|
||||
if ecosystem == "node":
|
||||
# Emergency-only: when the PATH node is missing/execution-probed broken
|
||||
# and the healthy bundled node was selected (skill-family precedence in
|
||||
# platform_layer), npm's `#!/usr/bin/env node` shebang must resolve the
|
||||
# working runtime, so the curated PATH gains the bundled-node dir. On
|
||||
# healthy systems the env stays byte-identical. npm itself is not
|
||||
# bundled: an absent npm still fails honestly upstream, and an npm
|
||||
# launcher with an ABSOLUTE node shebang ignoring PATH is a disclosed
|
||||
# residual.
|
||||
from ouroboros.platform_layer import skill_node_emergency_path_dir
|
||||
|
||||
prepend_dir = skill_node_emergency_path_dir()
|
||||
if prepend_dir:
|
||||
current = env.get("PATH", "")
|
||||
env["PATH"] = os.pathsep.join([prepend_dir, current]) if current else prepend_dir
|
||||
return env
|
||||
|
||||
|
||||
|
|
@ -178,6 +193,11 @@ def _install_python_packages(packages: List[str], env_root: pathlib.Path, timeou
|
|||
|
||||
|
||||
def _install_node_package(package: str, env_root: pathlib.Path, timeout_sec: int) -> List[Dict[str, Any]]:
|
||||
from ouroboros.platform_layer import bootstrap_process_path
|
||||
|
||||
# A GUI-launched macOS process starts with a truncated PATH; enrich it the
|
||||
# same way the process tools do BEFORE deciding npm is absent (T13).
|
||||
bootstrap_process_path()
|
||||
npm = shutil.which("npm")
|
||||
if not npm:
|
||||
raise RuntimeError("npm is not available on PATH")
|
||||
|
|
|
|||
230
ouroboros/node_runtime.py
Normal file
230
ouroboros/node_runtime.py
Normal file
|
|
@ -0,0 +1,230 @@
|
|||
"""Node runtime health and skill-family runtime policy.
|
||||
|
||||
Split out of ``platform_layer`` (which keeps the cross-platform primitives:
|
||||
``_hidden_run``, ``bootstrap_process_path``, ``resolve_bundled_node``): this
|
||||
module owns the EXECUTION-probed health verdicts and the skill-family
|
||||
precedence policy built on them. ``platform_layer`` re-exports the public
|
||||
names so existing importers keep working.
|
||||
|
||||
Platform access is by module attribute (``_platform.<name>``) on call, not by
|
||||
from-import, so the ``platform_layer -> node_runtime`` re-export cycle stays
|
||||
inert at import time.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import shutil
|
||||
import signal
|
||||
import subprocess
|
||||
from typing import Dict, List, NamedTuple, Tuple
|
||||
|
||||
from ouroboros import platform_layer as _platform
|
||||
|
||||
|
||||
def _probe_node_version_outcome(node_path: str, timeout_sec: float = 10) -> Tuple[str, str]:
|
||||
"""One ``node --version`` execution probe: ``(version, failure_reason)``.
|
||||
|
||||
Exactly one of the pair is non-empty. The reason vocabulary is small and
|
||||
typed-ish (``exec_failed:...``, ``timeout``, ``exit:N``, ``signal:NAME``)
|
||||
because the health consumers disclose it verbatim in traces and receipts.
|
||||
"""
|
||||
# A metadata probe must not inherit runtime/test hooks. In particular,
|
||||
# NODE_OPTIONS can contain test filters or preload modules that either make
|
||||
# `node --version` fail before the hermetic lane gets a chance to scrub the
|
||||
# variable or execute arbitrary operator code during a supposedly inert
|
||||
# version check.
|
||||
probe_env = dict(os.environ)
|
||||
probe_env.pop("NODE_OPTIONS", None)
|
||||
try:
|
||||
result = _platform._hidden_run(
|
||||
[str(node_path), "--version"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
encoding="utf-8",
|
||||
timeout=timeout_sec,
|
||||
check=False,
|
||||
env=probe_env,
|
||||
)
|
||||
except subprocess.TimeoutExpired:
|
||||
return "", "timeout"
|
||||
except (OSError, subprocess.SubprocessError) as exc:
|
||||
return "", f"exec_failed:{type(exc).__name__}"
|
||||
rc = int(result.returncode or 0)
|
||||
if rc != 0:
|
||||
if rc < 0:
|
||||
try:
|
||||
name = signal.Signals(abs(rc)).name
|
||||
except ValueError:
|
||||
name = f"SIG{abs(rc)}"
|
||||
return "", f"signal:{name}"
|
||||
return "", f"exit:{rc}"
|
||||
version = str(result.stdout or "").strip().removeprefix("v")
|
||||
if not version:
|
||||
# rc=0 with empty stdout: not a healthy node and not a named failure —
|
||||
# give the consumers a real reason instead of an empty pair.
|
||||
return "", "empty_version_output"
|
||||
return version, ""
|
||||
|
||||
|
||||
def probe_node_version(node_path: str) -> str:
|
||||
"""Return a normalized Node version, or ``""`` on probe failure."""
|
||||
version, _reason = _probe_node_version_outcome(node_path)
|
||||
return version
|
||||
|
||||
|
||||
class NodeRuntimeHealth(NamedTuple):
|
||||
"""Execution-probed usability of one Node executable path.
|
||||
|
||||
``status`` is ``healthy`` (probe returned a version), ``broken`` (the file
|
||||
exists but the probe failed — the incident class is a kernel
|
||||
SIGKILL/CODESIGNING death of a corrupted Homebrew node), or ``missing``.
|
||||
``reason`` carries the probe failure verbatim for traces/receipts.
|
||||
"""
|
||||
|
||||
status: str
|
||||
version: str = ""
|
||||
reason: str = ""
|
||||
path: str = ""
|
||||
# For a memoized ``timeout`` verdict: the probe budget it was observed
|
||||
# under. A caller with a LARGER budget re-probes instead of trusting it.
|
||||
probed_timeout: float = 0.0
|
||||
|
||||
@property
|
||||
def healthy(self) -> bool:
|
||||
return self.status == "healthy"
|
||||
|
||||
|
||||
# Process-local probe memo keyed by (lexical path, mtime_ns, size). A changed
|
||||
# binary re-probes; ``missing`` is deliberately never memoized so a runtime
|
||||
# installed mid-session (e.g. `brew reinstall node`, D9) is noticed without a
|
||||
# restart. A ``timeout`` verdict memoizes WITH the budget it was observed
|
||||
# under (``probed_timeout``): callers probe with DIFFERENT budgets (workspace
|
||||
# preflight uses a short 3s cap on the bounded admission path, skill surfaces
|
||||
# allow 10s), so a caller whose budget exceeds the cached one re-probes — a
|
||||
# short-budget timeout can therefore never poison a longer-budget consumer —
|
||||
# while a same-or-smaller budget reuses the verdict instead of stalling for
|
||||
# the full timeout again on every call (T15). The incident class (kernel
|
||||
# SIGKILL on launch) dies in milliseconds and memoizes normally. Known residual: a fix that does not
|
||||
# touch the file bytes (xattr / Gatekeeper requalification) keeps a stale
|
||||
# ``broken`` verdict for this process's lifetime — the trace disclosing the
|
||||
# memoized reason makes that visible rather than silent.
|
||||
_NODE_HEALTH_MEMO: Dict[Tuple[str, int, int], NodeRuntimeHealth] = {}
|
||||
|
||||
|
||||
def node_runtime_health(node_path: str, timeout_sec: float = 10) -> NodeRuntimeHealth:
|
||||
"""Probe (memoized) whether ``node_path`` is an actually-runnable Node.
|
||||
|
||||
``shutil.which`` proves only that a file exists and is executable; the
|
||||
incident class this exists for is a PATH node the kernel kills on launch,
|
||||
which only a real execution probe can see.
|
||||
"""
|
||||
text = str(node_path or "").strip()
|
||||
if not text:
|
||||
return NodeRuntimeHealth(status="missing", reason="empty_path")
|
||||
candidate = pathlib.Path(text)
|
||||
try:
|
||||
stat = candidate.stat()
|
||||
if not candidate.is_file() or not os.access(candidate, os.X_OK):
|
||||
return NodeRuntimeHealth(status="missing", reason="not_executable", path=text)
|
||||
except OSError as exc:
|
||||
return NodeRuntimeHealth(status="missing", reason=f"stat_failed:{type(exc).__name__}", path=text)
|
||||
key = (text, int(stat.st_mtime_ns), int(stat.st_size))
|
||||
cached = _NODE_HEALTH_MEMO.get(key)
|
||||
if cached is not None and (
|
||||
cached.reason != "timeout" or float(timeout_sec) <= cached.probed_timeout
|
||||
):
|
||||
return cached
|
||||
version, reason = _probe_node_version_outcome(text, timeout_sec=timeout_sec)
|
||||
if version:
|
||||
health = NodeRuntimeHealth(status="healthy", version=version, path=text)
|
||||
elif reason == "timeout":
|
||||
health = NodeRuntimeHealth(
|
||||
status="broken", reason=reason, path=text,
|
||||
probed_timeout=float(timeout_sec),
|
||||
)
|
||||
else:
|
||||
health = NodeRuntimeHealth(status="broken", reason=reason, path=text)
|
||||
_NODE_HEALTH_MEMO[key] = health
|
||||
return health
|
||||
|
||||
|
||||
def _path_node_runtime_health(timeout_sec: float = 10) -> NodeRuntimeHealth:
|
||||
"""Execution-probed health of the PATH-resolved ``node`` (post-bootstrap)."""
|
||||
_platform.bootstrap_process_path()
|
||||
located = shutil.which("node")
|
||||
if not located:
|
||||
return NodeRuntimeHealth(status="missing", reason="not_on_path")
|
||||
if not os.path.isabs(located):
|
||||
# A relative PATH entry resolves against THIS process's cwd here but
|
||||
# against the skill/companion cwd at exec time: health is unprovable,
|
||||
# so the rollback never trusts it (full-scope finding F-2, the same
|
||||
# contract as the generic resolver's T10 branch).
|
||||
return NodeRuntimeHealth(
|
||||
status="missing", reason="relative_path_entry_unprovable", path=located,
|
||||
)
|
||||
return node_runtime_health(located, timeout_sec=timeout_sec)
|
||||
|
||||
|
||||
def select_skill_node_runtime(timeout_sec: float = 10) -> Tuple[str, str]:
|
||||
"""The ONE owner of the skill-family Node runtime precedence.
|
||||
|
||||
Bundled-first: a skill's own runtime is the packaged, signed node when it
|
||||
is present AND execution-probed healthy (the python symmetry — python
|
||||
skills/companions already run on the embedded interpreter) — with a health
|
||||
ROLLBACK to a healthy PATH node when the bundled runtime is absent or
|
||||
provably broken. A candidate that fails its execution probe is never
|
||||
selected while a usable neighbour exists.
|
||||
|
||||
Returns ``(path, provenance)`` with provenance ``"bundled"`` or ``"path"``.
|
||||
When neither runtime is usable, returns ``("", reason)`` where the reason
|
||||
strings both probe verdicts verbatim for honest surface errors.
|
||||
"""
|
||||
facts: List[str] = []
|
||||
bundled = _platform.resolve_bundled_node()
|
||||
if bundled:
|
||||
bundled_health = node_runtime_health(bundled, timeout_sec=timeout_sec)
|
||||
if bundled_health.healthy:
|
||||
return bundled, "bundled"
|
||||
facts.append(f"bundled:{bundled_health.status}:{bundled_health.reason}")
|
||||
else:
|
||||
facts.append("bundled:absent")
|
||||
path_health = _path_node_runtime_health(timeout_sec=timeout_sec)
|
||||
if path_health.healthy:
|
||||
return path_health.path, "path"
|
||||
facts.append(
|
||||
f"path:{path_health.status}"
|
||||
+ (f":{path_health.reason}" if path_health.reason else "")
|
||||
)
|
||||
return "", "; ".join(facts)
|
||||
|
||||
|
||||
def skill_node_emergency_path_dir(timeout_sec: float = 10) -> str:
|
||||
"""Bundled-node dir to PREPEND to a skill-family child PATH — emergency only.
|
||||
|
||||
npm/npx/pnpm/yarn are NOT bundled and are never rewritten; their launchers
|
||||
resolve node via a ``#!/usr/bin/env node`` shebang. Only when the PATH node
|
||||
is missing or execution-probed broken AND the healthy bundled node was
|
||||
selected does the child PATH gain the bundled-node directory (on POSIX it
|
||||
contains exactly the ``node`` executable, so nothing else is shadowed; on
|
||||
Windows it is the node-standalone root holding ``node.exe``). On a healthy
|
||||
PATH this returns ``""`` and the child env stays byte-identical.
|
||||
|
||||
Known residuals (disclosed): an npm launcher rewritten to an absolute node
|
||||
shebang ignores PATH and keeps failing honestly; and the "nothing else is
|
||||
shadowed" claim holds for the official download scripts (which prune the
|
||||
archive to the bare node binary) — a custom ``OUROBOROS_BUNDLE_DIR``
|
||||
pointing at a FULL Node install would also shadow npm/npx with the bundled
|
||||
copies while the emergency is active.
|
||||
"""
|
||||
if not _platform.resolve_bundled_node():
|
||||
# No bundled runtime installed (source checkouts, bench clones): no
|
||||
# emergency lane exists and no probe is spent deciding that.
|
||||
return ""
|
||||
selected, provenance = select_skill_node_runtime(timeout_sec=timeout_sec)
|
||||
if not selected or provenance != "bundled":
|
||||
return ""
|
||||
if _path_node_runtime_health(timeout_sec=timeout_sec).healthy:
|
||||
return ""
|
||||
return str(pathlib.Path(selected).parent)
|
||||
|
|
@ -19,7 +19,12 @@ import uuid
|
|||
from dataclasses import dataclass, field
|
||||
from typing import Any, Callable, Dict, List, Optional, Tuple
|
||||
|
||||
from ouroboros.utils import atomic_write_json, replace_atomic, utc_now_iso
|
||||
from ouroboros.utils import (
|
||||
atomic_write_json,
|
||||
extract_trailing_json_object,
|
||||
replace_atomic,
|
||||
utc_now_iso,
|
||||
)
|
||||
|
||||
|
||||
OBSERVABILITY_DIR = "observability"
|
||||
|
|
@ -1221,13 +1226,42 @@ def latest_llm_response_text(drive_root: pathlib.Path, task_id: str) -> str:
|
|||
message = payload.get("message") if isinstance(payload, dict) else None
|
||||
content = message.get("content") if isinstance(message, dict) else None
|
||||
text = str(content or "").strip()
|
||||
if text and not _is_delivery_control_payload(text):
|
||||
return text
|
||||
if not text or _is_delivery_control_payload(text):
|
||||
continue
|
||||
# Lockstep with the loop's trailing-object parse: a body of prose
|
||||
# plus one TRAILING protocol object salvages the prose only — the
|
||||
# machine directive must never reach the owner's terminal result.
|
||||
# (A whole-body protocol object stays suppressed above; an object
|
||||
# with prose AFTER it is quoted material and passes through.)
|
||||
prose, parsed, duplicate_key = extract_trailing_json_object(
|
||||
text, duplicate_flag_keys=("delivery_control", "full_answer"),
|
||||
)
|
||||
if duplicate_key or (isinstance(parsed, dict) and "delivery_control" in parsed):
|
||||
if prose.strip():
|
||||
return prose.rstrip()
|
||||
continue
|
||||
return text
|
||||
except Exception:
|
||||
continue
|
||||
return ""
|
||||
|
||||
|
||||
def strip_protocol_fence(text: str) -> str:
|
||||
"""Strip ONE whole-body markdown fence, returning the trimmed inner body.
|
||||
|
||||
Shared normalization for every delivery-control protocol reader (the loop
|
||||
resolvers and the salvage predicate below): a fenced protocol object is
|
||||
still the protocol object. Anything short of a single fence spanning the
|
||||
whole body is returned stripped but otherwise unchanged.
|
||||
"""
|
||||
body = str(text or "").strip()
|
||||
if body.startswith("```"):
|
||||
first_break = body.find("\n")
|
||||
if first_break != -1 and body.endswith("```"):
|
||||
return body[first_break + 1:-3].strip()
|
||||
return body
|
||||
|
||||
|
||||
def _is_delivery_control_payload(text: str) -> bool:
|
||||
"""Whether persisted assistant text is the delivery-control PROTOCOL object.
|
||||
|
||||
|
|
@ -1237,14 +1271,13 @@ def _is_delivery_control_payload(text: str) -> bool:
|
|||
loop-side resolution used to let raw-salvage promote that JSON into the
|
||||
owner-facing terminal result. This is a STRUCTURAL typed-protocol check
|
||||
(exact JSON object carrying the protocol key, optionally in one markdown
|
||||
fence) — never semantic prose classification. Matching payloads stay
|
||||
forensic evidence in the observability store; they are simply not answers.
|
||||
fence) — never semantic prose classification: a MIXED prose+object answer
|
||||
stays salvageable here (this predicate has no latch knowledge; embedded-
|
||||
object containment belongs to the loop's latch-gated resolvers). Matching
|
||||
payloads stay forensic evidence in the observability store; they are
|
||||
simply not answers.
|
||||
"""
|
||||
body = str(text or "").strip()
|
||||
if body.startswith("```"):
|
||||
first_break = body.find("\n")
|
||||
if first_break != -1 and body.endswith("```"):
|
||||
body = body[first_break + 1:-3].strip()
|
||||
body = strip_protocol_fence(text)
|
||||
if not (body.startswith("{") and body.endswith("}")):
|
||||
return False
|
||||
try:
|
||||
|
|
|
|||
|
|
@ -172,6 +172,24 @@ REASON_REVIEW_CYCLES_EXHAUSTED = "review_cycles_exhausted"
|
|||
# objective terminalizes BLOCKED exactly like the spent-cap case above.
|
||||
REASON_REVIEW_QUORUM_UNREACHABLE = "plan_review_quorum_unreachable"
|
||||
|
||||
# A-material (owner ratification 2026-08-30): the agent resubmitted the SAME paid
|
||||
# acceptance identity — unchanged candidate answer AND no new obligation
|
||||
# disposition — so the recorded verdict was replayed for free instead of buying
|
||||
# another panel. The branch only fires when that recorded verdict was NOT a clean
|
||||
# PASS (a clean PASS replays as `accepted` on its own branch), so the objective
|
||||
# terminalizes BLOCKED exactly like the spent-cap case above: the deliverable was
|
||||
# never accepted and no further reviewer round will happen.
|
||||
REASON_IDENTICAL_ACCEPTANCE_REFUSED = "identical_acceptance_refused"
|
||||
|
||||
# The acceptance-decision reasons whose (finalized_unaccepted, reason) PAIR
|
||||
# terminalizes the objective axis BLOCKED. Value-keyed readers of the acceptance
|
||||
# decision live here: adding a terminal reason without adding it to the right key
|
||||
# is the silent false green DEVELOPMENT.md warns about.
|
||||
_ACCEPTANCE_BLOCKED_TERMINAL_REASONS = frozenset({
|
||||
REASON_REVIEW_CYCLES_EXHAUSTED,
|
||||
REASON_IDENTICAL_ACCEPTANCE_REFUSED,
|
||||
})
|
||||
|
||||
# CLOSED mapping: forced-finalization rail (the loop's typed reason_code) -> typed
|
||||
# acceptance-bypass reason, stamped by the loop's common forced-finalization recorder
|
||||
# when the panel was OWED (eligible) but a rail ended the task first. Both deadline
|
||||
|
|
@ -671,18 +689,22 @@ def _objective_axis(review: Dict[str, Any]) -> Dict[str, Any]:
|
|||
status = str(review.get("status") or "skipped")
|
||||
tier = str(review.get("outcome_tier") or "")
|
||||
decision = review.get("acceptance_decision") if isinstance(review.get("acceptance_decision"), dict) else {}
|
||||
_decision_reason = str(decision.get("reason") or "")
|
||||
if (
|
||||
str(decision.get("status") or "") == ACCEPTANCE_FINALIZED_UNACCEPTED
|
||||
and str(decision.get("reason") or "") == REASON_REVIEW_CYCLES_EXHAUSTED
|
||||
and _decision_reason in _ACCEPTANCE_BLOCKED_TERMINAL_REASONS
|
||||
):
|
||||
# D27: Required+Blocking acceptance whose shared cap is spent terminalizes
|
||||
# BLOCKED, whatever tier the last (failed) review proposed.
|
||||
# BLOCKED, whatever tier the last (failed) review proposed. A-material
|
||||
# (2026-08-30) adds the identical-paid-identity refusal on the same key:
|
||||
# the deliverable was never accepted and no further round will happen, so
|
||||
# the last review's proposed tier must not read as a green objective.
|
||||
return {
|
||||
"status": OBJECTIVE_FAIL,
|
||||
"source": "task_acceptance_review",
|
||||
"review_status": status,
|
||||
"outcome_tier": OUTCOME_TIER_BLOCKED,
|
||||
"reason": REASON_REVIEW_CYCLES_EXHAUSTED,
|
||||
"reason": _decision_reason,
|
||||
}
|
||||
if tier:
|
||||
# Reviewer tier is the canonical objective lexicon (completion-coach):
|
||||
|
|
|
|||
|
|
@ -23,6 +23,10 @@ KIND_FINALIZE_NOW = "finalize_now"
|
|||
# re-drain must restore the attempt latch; only terminal cleanup removes it).
|
||||
# Its ``text`` is the parser-required internal reason ("owner_hurry"), not prose.
|
||||
KIND_HURRY = "hurry"
|
||||
# Owner's verbatim quiz answer (#Q-2b): a typed control whose ``text`` is the
|
||||
# chosen option label (plus an optional owner comment) — delivered inside a
|
||||
# structural frame, never as forged free-form owner dialogue.
|
||||
KIND_QUIZ_ANSWER = "quiz_answer"
|
||||
# The mailbox is append-only, so a sender that changes its mind cannot delete the
|
||||
# control it already wrote — it appends this retraction naming the target msg_id.
|
||||
# Revocations are resolved by the READER over the whole mailbox, so a control that
|
||||
|
|
@ -194,7 +198,7 @@ def write_task_message(
|
|||
) -> bool:
|
||||
"""Write an addressed task-tree message without forging owner provenance."""
|
||||
|
||||
if provenance not in {"ancestor_task", "peer_via_ancestor", "system"}:
|
||||
if provenance not in {"ancestor_task", "peer_via_ancestor", "system", "descendant_task"}:
|
||||
return False
|
||||
path = _mailbox_path(drive_root, task_id)
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
|
@ -256,6 +260,10 @@ def deliver_task_message(
|
|||
prefix = f"[Message from task {relayed}, relayed by ancestor {source}]"
|
||||
elif provenance == "system":
|
||||
prefix = "[System task message]"
|
||||
elif provenance == "descendant_task":
|
||||
# Escalation direction is upward: signing it "ancestor" would invert
|
||||
# the sender's place in the tree (decision 31 hierarchy).
|
||||
prefix = f"[Escalation from descendant task {source}]"
|
||||
else:
|
||||
prefix = f"[Message from ancestor task {source}]"
|
||||
append_message(f"{prefix}\n{entry.get('text') or ''}")
|
||||
|
|
@ -269,6 +277,27 @@ def deliver_task_message(
|
|||
pass
|
||||
|
||||
|
||||
def deliver_quiz_answer(
|
||||
entry: Dict[str, Any], task_id: str, event_queue: Any, append_message: Any,
|
||||
) -> None:
|
||||
"""Inject a typed quiz answer (#Q-2b).
|
||||
|
||||
The entry's ``text`` is the complete host-authored frame (structural
|
||||
header + the owner's VERBATIM chosen label and optional comment, composed
|
||||
at ingress time where the projection block is in hand). The model judges
|
||||
freshness itself from the asked/answered timestamps in the frame — no
|
||||
host staleness verdict (owner decision 30=A)."""
|
||||
append_message(str(entry.get("text") or ""))
|
||||
if event_queue is not None:
|
||||
try:
|
||||
event_queue.put_nowait({
|
||||
"type": "quiz_answer_injected", "task_id": task_id,
|
||||
"msg_id": str(entry.get("msg_id") or ""),
|
||||
})
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def revoke_owner_control(
|
||||
drive_root: pathlib.Path, task_id: str, control_msg_id: str,
|
||||
) -> bool:
|
||||
|
|
|
|||
220
ouroboros/owner_quiz.py
Normal file
220
ouroboros/owner_quiz.py
Normal file
|
|
@ -0,0 +1,220 @@
|
|||
"""Durable owner-quiz lifecycle projection (#Q-2b, the answer half of the
|
||||
escalation channel).
|
||||
|
||||
This module owns everything quiz-lifecycle-specific — the asked/answered
|
||||
projection in the task result, the request-id idempotent answer write, and
|
||||
the terminal reconciliation — so the pinned surfaces (the escalate tool, the
|
||||
decision ingress, ``supervisor/queue_transitions``) keep thin dispatch only,
|
||||
mirroring ``ouroboros/owner_hurry.py``.
|
||||
|
||||
Shape inside ``task_results/<task_id>.json`` (canonical drive root, exactly
|
||||
like the hurry projection):
|
||||
|
||||
"owner_quiz": {
|
||||
"<quiz_id>": {
|
||||
"quiz_id", "question", "options": [label, ...], "stake",
|
||||
"assumption", "state": open|answered|expired_terminal,
|
||||
"asked_at", "answered_at"?, "answered_index"?, "request_id"?,
|
||||
"comment"?, "reconciled_at"?,
|
||||
}, ...
|
||||
}
|
||||
|
||||
Structural expiry only (owner decision 30=A): a quiz dies with its author —
|
||||
``reconcile_terminal`` runs on the task-done seam; there is no host TTL.
|
||||
The writer mutates ONLY the ``owner_quiz`` key via ``update_json_locked``
|
||||
(never ``write_task_result`` — its status-regression guard can drop the
|
||||
write), so concurrent terminal writers merge around it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, Callable, Dict, List, Optional
|
||||
|
||||
from ouroboros.utils import update_json_locked, utc_now_iso
|
||||
|
||||
STATE_OPEN = "open"
|
||||
STATE_ANSWERED = "answered"
|
||||
STATE_EXPIRED_TERMINAL = "expired_terminal"
|
||||
|
||||
# At most this many quiz blocks are retained per task (oldest evicted first);
|
||||
# a task asking more than this many questions keeps the recent ones live.
|
||||
_QUIZ_CAP = 16
|
||||
|
||||
|
||||
class _Keep:
|
||||
pass
|
||||
|
||||
|
||||
_KEEP = _Keep()
|
||||
|
||||
|
||||
def _quiz_result_path(drive_root: Any, task_id: str, *, create: bool = True):
|
||||
from ouroboros.task_results import task_result_path
|
||||
|
||||
return task_result_path(drive_root, str(task_id), create=create)
|
||||
|
||||
|
||||
def _mutate_projection(
|
||||
drive_root: Any, task_id: str,
|
||||
mutator: Callable[[Dict[str, Dict[str, Any]]], Any],
|
||||
) -> Dict[str, Dict[str, Any]]:
|
||||
"""Locked writer for the ``owner_quiz`` key only (hurry idiom).
|
||||
|
||||
``mutator(quizzes)`` mutates the dict in place and returns ``_KEEP`` to
|
||||
abort (file untouched) or anything else to commit. Returns the post-write
|
||||
(or pre-write, on abort) quizzes view.
|
||||
"""
|
||||
view: Dict[str, Dict[str, Any]] = {}
|
||||
|
||||
def _mutate(current: Dict[str, Any]) -> Optional[Dict[str, Any]]:
|
||||
view.clear()
|
||||
raw = current.get("owner_quiz")
|
||||
quizzes = {
|
||||
str(k): dict(v) for k, v in raw.items() if isinstance(v, dict)
|
||||
} if isinstance(raw, dict) else {}
|
||||
outcome = mutator(quizzes)
|
||||
if outcome is _KEEP:
|
||||
view.update(quizzes)
|
||||
return None
|
||||
if len(quizzes) > _QUIZ_CAP:
|
||||
# Evict CLOSED blocks first (oldest asked_at): an evicted OPEN
|
||||
# block would resurrect as an "Awaiting answer" card on replay
|
||||
# (the chat row froze state=open) whose click then 404s.
|
||||
def _eviction_key(key: str):
|
||||
block = quizzes[key]
|
||||
closed = str(block.get("state") or STATE_OPEN) != STATE_OPEN
|
||||
return (0 if closed else 1, str(block.get("asked_at") or ""))
|
||||
|
||||
for key in sorted(quizzes, key=_eviction_key)[:-_QUIZ_CAP]:
|
||||
quizzes.pop(key, None)
|
||||
updated = dict(current)
|
||||
updated["owner_quiz"] = quizzes
|
||||
view.update(quizzes)
|
||||
return updated
|
||||
|
||||
update_json_locked(_quiz_result_path(drive_root, task_id), _mutate)
|
||||
return view
|
||||
|
||||
|
||||
def record_asked(
|
||||
drive_root: Any, task_id: str, *,
|
||||
quiz_id: str, question: str, options: List[str],
|
||||
stake: str = "", assumption: str = "",
|
||||
) -> Dict[str, Any]:
|
||||
"""Worker-side projection write at ask time.
|
||||
|
||||
The stored option labels are the ingress's validation authority: an
|
||||
``option_index`` outside this list is refused, and the answer echoes the
|
||||
verbatim label back to the asking task."""
|
||||
stamp = utc_now_iso()
|
||||
block = {
|
||||
"quiz_id": str(quiz_id), "question": str(question or ""),
|
||||
"options": [str(label) for label in options],
|
||||
"stake": str(stake or ""), "assumption": str(assumption or ""),
|
||||
"state": STATE_OPEN, "asked_at": stamp,
|
||||
}
|
||||
|
||||
refused: Dict[str, str] = {}
|
||||
|
||||
def _mutator(quizzes: Dict[str, Dict[str, Any]]) -> Any:
|
||||
if str(quiz_id) in quizzes:
|
||||
return _KEEP # asked once; a redelivery never resets an answer
|
||||
open_count = sum(
|
||||
1 for row in quizzes.values()
|
||||
if str(row.get("state") or STATE_OPEN) == STATE_OPEN
|
||||
)
|
||||
if open_count >= _QUIZ_CAP:
|
||||
# Refuse the ask instead of letting the cap evict an OPEN block:
|
||||
# an evicted open card would stay clickable in chat but 404 on
|
||||
# answer. The asker already carries an assumption to proceed on.
|
||||
refused["reason"] = "open_quiz_cap"
|
||||
return _KEEP
|
||||
quizzes[str(quiz_id)] = block
|
||||
return True
|
||||
|
||||
view = _mutate_projection(drive_root, task_id, _mutator)
|
||||
if refused:
|
||||
return {"refused": refused["reason"]}
|
||||
return dict(view.get(str(quiz_id)) or block)
|
||||
|
||||
|
||||
def record_answered(
|
||||
drive_root: Any, task_id: str, *,
|
||||
quiz_id: str, option_index: int, request_id: str, comment: str = "",
|
||||
) -> Dict[str, Any]:
|
||||
"""Ingress-side answer write — request-id idempotent, first answer wins.
|
||||
|
||||
Returns ``{"ok", "state", "duplicate", "error", "block"}``:
|
||||
- unknown quiz_id → ``error="quiz_not_found"``;
|
||||
- open + valid index → answered (``ok=True``);
|
||||
- same ``request_id`` replay → the recorded confirmation, ``duplicate``;
|
||||
- already answered/expired with a different ``request_id`` → refusal with
|
||||
the truthful current ``state`` (the card settles, never re-invites);
|
||||
- out-of-range index → ``error="option_out_of_range"``.
|
||||
"""
|
||||
stamp = utc_now_iso()
|
||||
outcome: Dict[str, Any] = {}
|
||||
|
||||
def _mutator(quizzes: Dict[str, Dict[str, Any]]) -> Any:
|
||||
block = quizzes.get(str(quiz_id))
|
||||
if not isinstance(block, dict):
|
||||
outcome.update({"ok": False, "error": "quiz_not_found", "state": ""})
|
||||
return _KEEP
|
||||
state = str(block.get("state") or STATE_OPEN)
|
||||
if str(block.get("request_id") or "") and str(block.get("request_id")) == str(request_id or ""):
|
||||
outcome.update({"ok": True, "state": state, "duplicate": True, "block": dict(block)})
|
||||
return _KEEP
|
||||
if state != STATE_OPEN:
|
||||
outcome.update({"ok": False, "error": "quiz_closed", "state": state, "block": dict(block)})
|
||||
return _KEEP
|
||||
options = block.get("options") if isinstance(block.get("options"), list) else []
|
||||
if not isinstance(option_index, int) or not (0 <= option_index < len(options)):
|
||||
outcome.update({"ok": False, "error": "option_out_of_range", "state": state})
|
||||
return _KEEP
|
||||
block.update({
|
||||
"state": STATE_ANSWERED, "answered_at": stamp,
|
||||
"answered_index": int(option_index),
|
||||
"request_id": str(request_id or ""),
|
||||
**({"comment": str(comment)} if str(comment or "").strip() else {}),
|
||||
})
|
||||
quizzes[str(quiz_id)] = block
|
||||
outcome.update({"ok": True, "state": STATE_ANSWERED, "duplicate": False, "block": dict(block)})
|
||||
return True
|
||||
|
||||
_mutate_projection(drive_root, task_id, _mutator)
|
||||
return outcome
|
||||
|
||||
|
||||
def reconcile_terminal(drive_root: Any, task_id: str) -> List[str]:
|
||||
"""Task-done reconciliation: every still-open quiz expires structurally.
|
||||
|
||||
Returns the quiz ids that flipped to ``expired_terminal`` so the caller
|
||||
can emit their ``quiz_state`` frames. Never resurrects or rewrites an
|
||||
answered block."""
|
||||
stamp = utc_now_iso()
|
||||
expired: List[str] = []
|
||||
|
||||
def _mutator(quizzes: Dict[str, Dict[str, Any]]) -> Any:
|
||||
for key, block in quizzes.items():
|
||||
if str(block.get("state") or STATE_OPEN) == STATE_OPEN:
|
||||
block.update({"state": STATE_EXPIRED_TERMINAL, "reconciled_at": stamp})
|
||||
expired.append(str(key))
|
||||
return True if expired else _KEEP
|
||||
|
||||
_mutate_projection(drive_root, task_id, _mutator)
|
||||
return expired
|
||||
|
||||
|
||||
def quiz_states(drive_root: Any, task_id: str) -> Dict[str, Dict[str, Any]]:
|
||||
"""Read-only view of the projection for history replay merge."""
|
||||
try:
|
||||
import json
|
||||
|
||||
# create=False: a read-only replay view must not mkdir as a side effect.
|
||||
raw = json.loads(_quiz_result_path(drive_root, task_id, create=False).read_text(encoding="utf-8"))
|
||||
except Exception:
|
||||
return {}
|
||||
quizzes = raw.get("owner_quiz") if isinstance(raw, dict) else None
|
||||
if not isinstance(quizzes, dict):
|
||||
return {}
|
||||
return {str(k): dict(v) for k, v in quizzes.items() if isinstance(v, dict)}
|
||||
|
|
@ -1038,32 +1038,6 @@ def node_distribution_platform() -> str:
|
|||
return ""
|
||||
|
||||
|
||||
def probe_node_version(node_path: str) -> str:
|
||||
"""Return a normalized bundled-Node version, or ``""`` on probe failure."""
|
||||
# A metadata probe must not inherit runtime/test hooks. In particular,
|
||||
# NODE_OPTIONS can contain test filters or preload modules that either make
|
||||
# `node --version` fail before the hermetic lane gets a chance to scrub the
|
||||
# variable or execute arbitrary operator code during a supposedly inert
|
||||
# version check.
|
||||
probe_env = dict(os.environ)
|
||||
probe_env.pop("NODE_OPTIONS", None)
|
||||
try:
|
||||
result = _hidden_run(
|
||||
[str(node_path), "--version"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
encoding="utf-8",
|
||||
timeout=10,
|
||||
check=False,
|
||||
env=probe_env,
|
||||
)
|
||||
except (OSError, subprocess.SubprocessError):
|
||||
return ""
|
||||
if result.returncode != 0:
|
||||
return ""
|
||||
return str(result.stdout or "").strip().removeprefix("v")
|
||||
|
||||
|
||||
def embedded_ripgrep_candidates(base_dir: pathlib.Path) -> List[pathlib.Path]:
|
||||
"""Return candidate bundled ripgrep paths."""
|
||||
if IS_WINDOWS:
|
||||
|
|
@ -1449,3 +1423,30 @@ def resume_process(pid: int) -> bool:
|
|||
except Exception as exc:
|
||||
log.warning("resume_process failed: %s", exc)
|
||||
return False
|
||||
|
||||
|
||||
# Node runtime health/policy moved to ouroboros/node_runtime.py (its own module:
|
||||
# the policy grew past what a cross-platform primitives file should hold, and
|
||||
# the 1600-line module gate agrees). The re-export is a PEP 562 module
|
||||
# __getattr__ rather than an eager from-import: node_runtime itself imports
|
||||
# this module at module level, and an eager import back from HERE re-entered a
|
||||
# partially initialized node_runtime whenever node_runtime was imported first
|
||||
# (triad finding, all three phase-C reviewers). Lazy resolution keeps both
|
||||
# import orders sound while every existing importer (preflight_node, skill
|
||||
# surfaces, the interpreter resolver, claudexor_runtime) keeps its
|
||||
# `from ouroboros.platform_layer import <name>` spelling unchanged.
|
||||
_NODE_RUNTIME_REEXPORTS = (
|
||||
"NodeRuntimeHealth",
|
||||
"node_runtime_health",
|
||||
"probe_node_version",
|
||||
"select_skill_node_runtime",
|
||||
"skill_node_emergency_path_dir",
|
||||
)
|
||||
|
||||
|
||||
def __getattr__(name: str):
|
||||
if name in _NODE_RUNTIME_REEXPORTS:
|
||||
from ouroboros import node_runtime
|
||||
|
||||
return getattr(node_runtime, name)
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
|
|
|
|||
920
ouroboros/process_interpreters.py
Normal file
920
ouroboros/process_interpreters.py
Normal file
|
|
@ -0,0 +1,920 @@
|
|||
"""Surface-aware interpreter/runtime selection for host process tools.
|
||||
|
||||
Only the four public process launch surfaces opt into these resolvers, and
|
||||
the two families run at DIFFERENT points of the dispatch pipeline: Python
|
||||
resolves pre-dispatch (never executes a candidate, so guards and handler see
|
||||
the same argv), while Node resolves post-gates (its health probe EXECUTES the
|
||||
candidate — see ``resolve_process_node``), so the deterministic guards always
|
||||
inspect the original bare argv and the handler sees the attested rewrite.
|
||||
Launchers must not rewrite the interpreter afterwards.
|
||||
|
||||
Two families share one trace/attestation contract:
|
||||
|
||||
* Python (``resolve_process_python``): the original five-step ladder ending in
|
||||
a typed pre-block when no interpreter can be proven.
|
||||
* Node (``resolve_process_node``): PATH-first with an execution health probe
|
||||
(``shutil.which`` proves only that a file exists — the incident class is a
|
||||
PATH node the kernel SIGKILLs on launch), falling back to the bundled signed
|
||||
runtime only when the PATH candidate is missing or probe-dead. Node never
|
||||
pre-blocks: with no usable runtime the argv runs as written and fails
|
||||
honestly, while the trace discloses the probe facts.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import os
|
||||
import pathlib
|
||||
import shutil
|
||||
import sys
|
||||
from dataclasses import asdict, dataclass
|
||||
from typing import Any, Dict, Mapping, Optional
|
||||
|
||||
from ouroboros.contracts.task_constraint import (
|
||||
TaskConstraint,
|
||||
normalize_task_constraint,
|
||||
)
|
||||
from ouroboros.platform_layer import (
|
||||
IS_WINDOWS,
|
||||
PATH_SEP,
|
||||
bootstrap_process_path,
|
||||
node_runtime_health,
|
||||
project_venv_python,
|
||||
resolve_bundled_node,
|
||||
)
|
||||
from ouroboros.shell_parse import normalize_check_argv, shell_command_string, shell_tokens
|
||||
from ouroboros.tool_access import (
|
||||
ResolvedResourceBinding,
|
||||
build_resolved_resource_binding,
|
||||
path_is_relative_to,
|
||||
)
|
||||
from ouroboros.utils import append_jsonl, utc_now_iso
|
||||
|
||||
_PYTHON_TOKENS = frozenset({"python", "python3"})
|
||||
# argv[0] spellings the node resolver may REWRITE to the bundled runtime.
|
||||
_NODE_TOKENS = frozenset({"node", "nodejs"})
|
||||
# The wider node family that TRIGGERS the emergency PATH prepend (npm & co are
|
||||
# not shipped in the bundle, so they are never rewritten — their env shebang
|
||||
# `#!/usr/bin/env node` picks up the prepended runtime instead; a formula with
|
||||
# a rewritten absolute shebang is a disclosed residual, Q2-1=A).
|
||||
_NODE_FAMILY_TOKENS = frozenset({"node", "nodejs", "npm", "npx", "pnpm", "yarn", "corepack"})
|
||||
# Shell wrappers whose ``-c`` body is seam-scanned for family tokens (R3).
|
||||
_NODE_SHELL_WRAPPERS = frozenset({"sh", "bash", "zsh", "dash"})
|
||||
_WINDOWS_LAUNCHER_SUFFIXES = (".exe", ".cmd", ".bat")
|
||||
_PROCESS_TOOLS = frozenset({"run_command", "run_script", "start_service", "verify_and_record"})
|
||||
# Mirrors the run-kind set of tools/verify.py (`_RUN_KINDS`); keep in sync when
|
||||
# a new verify run-kind is introduced there.
|
||||
_VERIFY_RUN_KINDS = frozenset({"visible_verifier", "explicit_command", "explicit_metric"})
|
||||
_VERIFIED_RESOLUTIONS = frozenset(
|
||||
{
|
||||
("reviewed_skill_environment", "isolated_skill"),
|
||||
("executor_backend_python3", "backend_path"),
|
||||
("project_venv", "project_venv"),
|
||||
("agent_python", "ouroboros_agent"),
|
||||
# Node resolutions proved by an execution probe (or delegated to a
|
||||
# non-local backend, mirroring executor_backend_python3).
|
||||
("executor_backend_node", "backend_path"),
|
||||
("path_node_healthy", "host_path"),
|
||||
("bundled_node_fallback", "bundled_node"),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class InterpreterResolutionTrace:
|
||||
"""One resolver decision for one process-tool call, any family.
|
||||
|
||||
``resolved_interpreter`` is the EXECUTION identity: for a no-op resolution
|
||||
it equals ``requested_interpreter`` (argv stays byte-identical); when the
|
||||
resolver substitutes a runtime it is the absolute path execution uses. For
|
||||
``verify_and_record`` the node resolver never rewrites ``args["check"]``
|
||||
(the receipt keeps the original text as its identity, amendment R4) — the
|
||||
handler reads the substitution from this attestation instead.
|
||||
"""
|
||||
|
||||
tool: str
|
||||
requested_interpreter: str
|
||||
resolved_interpreter: str
|
||||
surface: str
|
||||
environment: str
|
||||
reason: str
|
||||
fallback_reason: str = ""
|
||||
error_reason: str = ""
|
||||
target_root: str = ""
|
||||
target_cwd: str = ""
|
||||
target_source: str = ""
|
||||
target_skill: str = ""
|
||||
family: str = "python"
|
||||
# Node-family provenance: the frozen PATH the resolver probed (and the base
|
||||
# of any attested child-env prepend), the physically identified executable,
|
||||
# its probed version, and the emergency bundled-runtime dir to prepend.
|
||||
path_snapshot: str = ""
|
||||
env_path_prepend: str = ""
|
||||
runtime_path: str = ""
|
||||
runtime_version: str = ""
|
||||
|
||||
@property
|
||||
def changed(self) -> bool:
|
||||
return self.requested_interpreter != self.resolved_interpreter
|
||||
|
||||
@property
|
||||
def verified(self) -> bool:
|
||||
"""Whether the resolver proved the selected interpreter provenance."""
|
||||
|
||||
return (self.reason, self.environment) in _VERIFIED_RESOLUTIONS
|
||||
|
||||
def to_event(self) -> Dict[str, Any]:
|
||||
event = {**asdict(self), "changed": self.changed}
|
||||
if self.family == "python":
|
||||
# The long-standing python event payload stays byte-identical; the
|
||||
# generalization fields ride only on non-python families.
|
||||
for key in ("family", "path_snapshot", "env_path_prepend", "runtime_path", "runtime_version"):
|
||||
event.pop(key, None)
|
||||
return event
|
||||
|
||||
|
||||
# Existing isinstance-consumers and tests keep working under the historic name.
|
||||
PythonResolutionTrace = InterpreterResolutionTrace
|
||||
|
||||
|
||||
def _python_request(tool_name: str, args: Mapping[str, Any]) -> tuple[str, list[str] | None]:
|
||||
"""Return the exact eligible token and normalized argv, when applicable."""
|
||||
|
||||
if tool_name in {"run_command", "start_service"}:
|
||||
raw = args.get("cmd")
|
||||
if not isinstance(raw, list) or not raw:
|
||||
return "", None
|
||||
argv = [str(part) for part in raw]
|
||||
requested = str(argv[0]).strip()
|
||||
return (requested, argv) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
if tool_name == "run_script":
|
||||
requested = str(args.get("interpreter") or "python3").strip() or "python3"
|
||||
return (requested, None) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
if tool_name == "verify_and_record":
|
||||
kind = str(args.get("contract_kind") or "").strip()
|
||||
if kind not in _VERIFY_RUN_KINDS:
|
||||
return "", None
|
||||
argv = normalize_check_argv(args.get("check")) or []
|
||||
if not argv:
|
||||
return "", None
|
||||
requested = str(argv[0]).strip()
|
||||
return (requested, argv) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
return "", None
|
||||
|
||||
|
||||
def _usable_executable(path_text: str) -> str:
|
||||
"""Validate an interpreter while preserving venv symlink semantics."""
|
||||
|
||||
text = str(path_text or "").strip()
|
||||
if not text:
|
||||
return ""
|
||||
candidate = pathlib.Path(text).expanduser()
|
||||
if not candidate.is_absolute():
|
||||
located = shutil.which(text)
|
||||
if not located:
|
||||
return ""
|
||||
candidate = pathlib.Path(located)
|
||||
try:
|
||||
if not candidate.is_file() or not os.access(candidate, os.X_OK):
|
||||
return ""
|
||||
except OSError:
|
||||
return ""
|
||||
# Do not resolve a venv's python symlink: executing the lexical path is what
|
||||
# lets Python discover the adjacent pyvenv.cfg and preserve the environment.
|
||||
return os.path.abspath(os.fspath(candidate))
|
||||
|
||||
|
||||
def _reviewed_skill_python(
|
||||
ctx: Any,
|
||||
binding: ResolvedResourceBinding | None = None,
|
||||
) -> tuple[str, str]:
|
||||
"""Return the lifecycle-proven isolated Python for the selected skill.
|
||||
|
||||
A dispatch binding is authoritative and loads exactly its physical payload
|
||||
against its canonical state root. Legacy task metadata is consulted only
|
||||
when a non-registry/direct caller supplied no binding.
|
||||
"""
|
||||
|
||||
try:
|
||||
from ouroboros.marketplace.isolated_deps import python_runtime_binary, read_deps_state
|
||||
from ouroboros.skill_loader import find_skill, load_skill
|
||||
from ouroboros.skill_readiness import skill_readiness_for_execution
|
||||
|
||||
if binding is not None:
|
||||
if binding.root != "skill_payload":
|
||||
return "", ""
|
||||
drive_root = pathlib.Path(binding.state_drive_root)
|
||||
loaded = load_skill(pathlib.Path(binding.base_path), drive_root)
|
||||
if loaded is None or loaded.name != binding.skill_name:
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
else:
|
||||
metadata = getattr(ctx, "task_metadata", {})
|
||||
metadata = metadata if isinstance(metadata, dict) else {}
|
||||
skill_name = str(metadata.get("skill") or "").strip()
|
||||
if not skill_name:
|
||||
return "", ""
|
||||
drive_root = pathlib.Path(getattr(ctx, "drive_root"))
|
||||
loaded = find_skill(drive_root, skill_name)
|
||||
if loaded is None or not skill_readiness_for_execution(drive_root, loaded).ready:
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
deps_state = read_deps_state(drive_root, loaded.name, loaded.skill_dir)
|
||||
if str(deps_state.get("status") or "") != "installed":
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
candidate = python_runtime_binary(loaded.skill_dir)
|
||||
usable = _usable_executable(str(candidate or ""))
|
||||
if usable:
|
||||
return usable, ""
|
||||
except Exception:
|
||||
return "", "reviewed_skill_environment_probe_failed"
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
|
||||
|
||||
def _executor_covers_kind(ctx: Any, work_dir: pathlib.Path) -> tuple[bool, str, str]:
|
||||
"""Whether a configured executor covers ``work_dir``, plus its backend kind."""
|
||||
|
||||
try:
|
||||
from ouroboros.workspace_executor import executor_ref_from_ctx, map_host_path
|
||||
|
||||
executor = executor_ref_from_ctx(ctx)
|
||||
if executor is None:
|
||||
return False, "", ""
|
||||
map_host_path(executor, pathlib.Path(work_dir).resolve(strict=False))
|
||||
return True, str(executor.kind or ""), ""
|
||||
except ValueError:
|
||||
return False, "", ""
|
||||
except Exception:
|
||||
return False, "", "executor_resolution_failed"
|
||||
|
||||
|
||||
def _executor_covers(ctx: Any, work_dir: pathlib.Path) -> tuple[bool, str]:
|
||||
covers, _kind, error = _executor_covers_kind(ctx, work_dir)
|
||||
return covers, error
|
||||
|
||||
|
||||
def _surface_for(
|
||||
ctx: Any,
|
||||
binding: ResolvedResourceBinding,
|
||||
constraint: Optional[TaskConstraint],
|
||||
) -> str:
|
||||
if binding.root != "active_workspace":
|
||||
return binding.root or "unresolved"
|
||||
if constraint and constraint.mode == "acting_subagent" and constraint.surface == "self_worktree":
|
||||
return "system_repo"
|
||||
mode = str(getattr(ctx, "workspace_mode", "") or "").strip().lower()
|
||||
if mode in {"external", "external_workspace", "genesis"}:
|
||||
return "external_workspace"
|
||||
system_repo = pathlib.Path(
|
||||
getattr(ctx, "system_repo_dir", None) or getattr(ctx, "repo_dir", binding.target_path)
|
||||
).resolve(strict=False)
|
||||
return "system_repo" if path_is_relative_to(binding.target_path, system_repo) else "external_workspace"
|
||||
|
||||
|
||||
def _trace_target(binding: ResolvedResourceBinding) -> Dict[str, str]:
|
||||
return {
|
||||
"target_root": binding.root,
|
||||
"target_cwd": str(binding.target_path),
|
||||
"target_source": binding.source,
|
||||
"target_skill": binding.skill_name,
|
||||
}
|
||||
|
||||
|
||||
def _project_root(ctx: Any, surface: str, work_dir: pathlib.Path) -> pathlib.Path:
|
||||
if surface == "external_workspace":
|
||||
workspace_root = getattr(ctx, "workspace_root", None)
|
||||
if workspace_root:
|
||||
candidate = pathlib.Path(workspace_root).resolve(strict=False)
|
||||
if path_is_relative_to(work_dir, candidate):
|
||||
return candidate
|
||||
return work_dir
|
||||
|
||||
|
||||
def _replace_request(
|
||||
tool_name: str,
|
||||
args: Mapping[str, Any],
|
||||
argv: list[str] | None,
|
||||
resolved: str,
|
||||
) -> Dict[str, Any]:
|
||||
out = dict(args)
|
||||
if tool_name in {"run_command", "start_service"}:
|
||||
new_argv = list(argv or [])
|
||||
new_argv[0] = resolved
|
||||
out["cmd"] = new_argv
|
||||
elif tool_name == "run_script":
|
||||
out["interpreter"] = resolved
|
||||
elif tool_name == "verify_and_record":
|
||||
new_argv = list(argv or [])
|
||||
new_argv[0] = resolved
|
||||
out["check"] = new_argv
|
||||
return out
|
||||
|
||||
|
||||
def resolve_process_python(
|
||||
ctx: Any,
|
||||
tool_name: str,
|
||||
args: Mapping[str, Any],
|
||||
*,
|
||||
runtime_mode: str,
|
||||
effective_constraint: Optional[TaskConstraint] = None,
|
||||
resolved_binding: ResolvedResourceBinding | None = None,
|
||||
) -> tuple[Dict[str, Any], Optional[PythonResolutionTrace]]:
|
||||
"""Resolve an exact ``python``/``python3`` request for one process tool."""
|
||||
|
||||
name = str(tool_name or "").strip()
|
||||
original = dict(args or {})
|
||||
if name not in _PROCESS_TOOLS:
|
||||
return original, None
|
||||
requested, argv = _python_request(name, original)
|
||||
if not requested:
|
||||
return original, None
|
||||
|
||||
constraint = normalize_task_constraint(effective_constraint)
|
||||
cwd_text = str(original.get("cwd") or "")
|
||||
binding = resolved_binding
|
||||
try:
|
||||
operation = "service" if name == "start_service" else "shell"
|
||||
if binding is None:
|
||||
binding = build_resolved_resource_binding(
|
||||
ctx,
|
||||
operation=operation,
|
||||
process_cwd=cwd_text,
|
||||
bucket=str(original.get("bucket") or ""),
|
||||
skill_name=str(original.get("skill_name") or ""),
|
||||
)
|
||||
work_dir = pathlib.Path(binding.target_path).resolve(strict=False)
|
||||
except Exception:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface="unresolved",
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason="cwd_resolution_failed",
|
||||
error_reason="cwd_resolution_failed",
|
||||
)
|
||||
return original, trace
|
||||
|
||||
fallback_reason = ""
|
||||
assert binding is not None
|
||||
skill_binding = resolved_binding
|
||||
if skill_binding is None and binding.root == "skill_payload":
|
||||
skill_binding = binding
|
||||
skill_python, skill_reason = _reviewed_skill_python(ctx, skill_binding)
|
||||
if skill_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=skill_python,
|
||||
surface="reviewed_skill",
|
||||
environment="isolated_skill",
|
||||
reason="reviewed_skill_environment",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, skill_python), trace
|
||||
if skill_reason:
|
||||
fallback_reason = skill_reason
|
||||
|
||||
executor_active, executor_error = _executor_covers(ctx, work_dir)
|
||||
if executor_active:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter="python3",
|
||||
surface="executor",
|
||||
environment="backend_path",
|
||||
reason="executor_backend_python3",
|
||||
fallback_reason=fallback_reason,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, "python3"), trace
|
||||
if executor_error and not fallback_reason:
|
||||
fallback_reason = executor_error
|
||||
|
||||
surface = _surface_for(ctx, binding, constraint)
|
||||
if surface in {"external_workspace", "user_files"}:
|
||||
project_python = project_venv_python(_project_root(ctx, surface, work_dir))
|
||||
if project_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=project_python,
|
||||
surface=surface,
|
||||
environment="project_venv",
|
||||
reason="project_venv",
|
||||
fallback_reason=fallback_reason,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, project_python), trace
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason=fallback_reason or "project_venv_unavailable",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
|
||||
configured_agent_python = _usable_executable(
|
||||
os.environ.get("OUROBOROS_AGENT_PYTHON", "")
|
||||
)
|
||||
agent_python = configured_agent_python or _usable_executable(sys.executable or "")
|
||||
if agent_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=agent_python,
|
||||
surface=surface,
|
||||
environment="ouroboros_agent",
|
||||
reason="agent_python",
|
||||
fallback_reason=(
|
||||
fallback_reason
|
||||
or ("agent_env_unavailable_process_fallback" if not configured_agent_python else "")
|
||||
),
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, agent_python), trace
|
||||
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason=fallback_reason or "agent_python_unavailable",
|
||||
error_reason="agent_python_unavailable",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
|
||||
|
||||
def _normalize_runtime_token(text: str) -> str:
|
||||
"""Trimmed token for family matching; case-folding and launcher-suffix
|
||||
stripping apply ONLY on Windows (R7/T9): there ``node.exe``/``NPM.CMD``
|
||||
classify like their bare spellings, while POSIX exec is case-sensitive so
|
||||
the exact token is preserved. An absolute path or a versioned name
|
||||
(``node20``) never equals a family token, so both stay untouched by
|
||||
construction — same contract as python.
|
||||
"""
|
||||
|
||||
token = str(text or "").strip()
|
||||
if not IS_WINDOWS:
|
||||
# POSIX exec is case-sensitive: "NODE" is a different file, and
|
||||
# launcher suffixes are a Windows-only convention.
|
||||
return token
|
||||
token = token.lower()
|
||||
for suffix in _WINDOWS_LAUNCHER_SUFFIXES:
|
||||
if token.endswith(suffix):
|
||||
return token[: -len(suffix)]
|
||||
return token
|
||||
|
||||
|
||||
def _shell_body_names_node_family(body: str) -> bool:
|
||||
"""Deterministic seam-scan of a shell body for node-family tokens (R3).
|
||||
|
||||
Reuses the guard-layer tokenizer (``shell_parse.shell_tokens``) — no regex
|
||||
over prose. A transitive spawn (a python script that itself execs node)
|
||||
stays a documented residual.
|
||||
"""
|
||||
|
||||
tokens = shell_tokens(str(body or "")) or []
|
||||
return any(_normalize_runtime_token(token) in _NODE_FAMILY_TOKENS for token in tokens)
|
||||
|
||||
|
||||
def _node_request(tool_name: str, args: Mapping[str, Any]) -> tuple[str, list[str] | None, str]:
|
||||
"""Return ``(requested_token, argv, trigger)`` for a node-family launch.
|
||||
|
||||
``trigger`` is ``""`` (not a node-family request), ``"runtime"`` (argv[0]
|
||||
is node/nodejs — rewrite-eligible), ``"family"`` (npm/npx/pnpm/yarn/
|
||||
corepack — emergency prepend only), or ``"shell_body"`` (a shell wrapper —
|
||||
sh/bash/zsh/dash, matched by basename — whose command body names a family
|
||||
tool; emergency prepend only).
|
||||
"""
|
||||
|
||||
requested = ""
|
||||
argv: list[str] | None = None
|
||||
body = ""
|
||||
if tool_name in {"run_command", "start_service"}:
|
||||
raw = args.get("cmd")
|
||||
if not isinstance(raw, list) or not raw:
|
||||
return "", None, ""
|
||||
argv = [str(part) for part in raw]
|
||||
requested = argv[0].strip()
|
||||
if requested != argv[0]:
|
||||
# A whitespace-padded head is not a bare runtime token: classifying
|
||||
# it would let the trace attest a substitution that deferred
|
||||
# executors (verify) never apply — run it as written instead (T8).
|
||||
return "", None, ""
|
||||
body = shell_command_string(argv)
|
||||
elif tool_name == "run_script":
|
||||
requested = str(args.get("interpreter") or "python3").strip() or "python3"
|
||||
body = str(args.get("script") or "")
|
||||
elif tool_name == "verify_and_record":
|
||||
kind = str(args.get("contract_kind") or "").strip()
|
||||
if kind not in _VERIFY_RUN_KINDS:
|
||||
return "", None, ""
|
||||
argv = normalize_check_argv(args.get("check")) or []
|
||||
if not argv:
|
||||
return "", None, ""
|
||||
requested = str(argv[0]).strip()
|
||||
if requested != str(argv[0]):
|
||||
return "", None, "" # padded head: not bare, run as written (T8)
|
||||
body = shell_command_string(argv)
|
||||
else:
|
||||
return "", None, ""
|
||||
|
||||
normalized = _normalize_runtime_token(requested)
|
||||
if normalized in _NODE_TOKENS:
|
||||
return requested, argv, "runtime"
|
||||
if normalized in _NODE_FAMILY_TOKENS:
|
||||
return requested, argv, "family"
|
||||
# Wrapper matching is BASENAME-based (full-scope finding F-1): /bin/sh and
|
||||
# zsh -c bodies reproduce the incident class exactly like bare sh. This
|
||||
# grants no new exec power — a wrapper hit only scans the body and rides
|
||||
# the env prepend; family/runtime tokens above stay bare-only on purpose
|
||||
# (an absolute node path is the caller's explicit runtime choice).
|
||||
wrapper_token = _normalize_runtime_token(pathlib.PurePath(requested).name)
|
||||
if wrapper_token in _NODE_SHELL_WRAPPERS and _shell_body_names_node_family(body):
|
||||
return requested, argv, "shell_body"
|
||||
return "", None, ""
|
||||
|
||||
|
||||
def _node_trace(**fields: Any) -> InterpreterResolutionTrace:
|
||||
return InterpreterResolutionTrace(family="node", **fields)
|
||||
|
||||
|
||||
def resolve_process_node(
|
||||
ctx: Any,
|
||||
tool_name: str,
|
||||
args: Mapping[str, Any],
|
||||
*,
|
||||
runtime_mode: str,
|
||||
effective_constraint: Optional[TaskConstraint] = None,
|
||||
resolved_binding: ResolvedResourceBinding | None = None,
|
||||
) -> tuple[Dict[str, Any], Optional[InterpreterResolutionTrace]]:
|
||||
"""Resolve a node-family request for one process tool (D2/D3/D4 ladder).
|
||||
|
||||
Ladder: (i) a NON-local executor backend resolves ``node`` in its own
|
||||
filesystem — argv untouched, no host path leaks into the container; a
|
||||
LOCAL executor runs on this host, so the ladder continues (Q2-3). (ii) a
|
||||
PATH candidate that passes the execution health probe wins — argv and env
|
||||
stay byte-identical (the healthy-system no-op invariant). (iii) a missing
|
||||
or probe-dead PATH candidate falls back to the healthy bundled runtime:
|
||||
node/nodejs argv[0] is rewritten, and EVERY triggered launch gets the
|
||||
bundled dir attested as a child-env PATH prepend (fixes npm & co and
|
||||
``sh -c`` bodies). (iv) with neither usable the argv runs as written and
|
||||
fails honestly — node never raises a typed pre-block (R8).
|
||||
|
||||
Placement (post-gates, deliberate — differs from python's pre-guard seam):
|
||||
the ladder's health check is an EXECUTION probe (``<candidate> --version``)
|
||||
and argv[0] steers which executable it runs. Pre-guard it would execute an
|
||||
agent-influenced binary BEFORE the light fence / protected-write gates /
|
||||
shell guard / safety supervisor — a planted PATH shim named ``node`` would
|
||||
run its payload on a call those gates then refuse. Post-gates the probe
|
||||
holds strictly LESS power than the just-approved call itself. Guards
|
||||
therefore inspect the ORIGINAL bare argv; the substitution the handler
|
||||
executes is limited to a family-stable argv[0]
|
||||
(``shell_guards.interpreter_family`` classifies the bare token and the
|
||||
resolver's absolute path identically) and is disclosed via the trace event
|
||||
and the per-call attestation. Node never pre-blocks (R8): with no usable
|
||||
runtime the argv runs as written and fails honestly.
|
||||
"""
|
||||
|
||||
name = str(tool_name or "").strip()
|
||||
original = dict(args or {})
|
||||
if name not in _PROCESS_TOOLS:
|
||||
return original, None
|
||||
requested, argv, trigger = _node_request(name, original)
|
||||
if not trigger:
|
||||
return original, None
|
||||
|
||||
# R1: the resolver must see the SAME PATH the handler will (both call the
|
||||
# idempotent bootstrap; whoever runs first performs the one mutation). The
|
||||
# snapshot is the base of any attested child-env PATH prepend.
|
||||
bootstrap_process_path()
|
||||
path_snapshot = str(os.environ.get("PATH", "") or "")
|
||||
|
||||
constraint = normalize_task_constraint(effective_constraint)
|
||||
cwd_text = str(original.get("cwd") or "")
|
||||
binding = resolved_binding
|
||||
try:
|
||||
operation = "service" if name == "start_service" else "shell"
|
||||
if binding is None:
|
||||
binding = build_resolved_resource_binding(
|
||||
ctx,
|
||||
operation=operation,
|
||||
process_cwd=cwd_text,
|
||||
bucket=str(original.get("bucket") or ""),
|
||||
skill_name=str(original.get("skill_name") or ""),
|
||||
)
|
||||
work_dir = pathlib.Path(binding.target_path).resolve(strict=False)
|
||||
except Exception:
|
||||
# No typed pre-block for node (R8): execution proceeds as written and
|
||||
# the handler's own binding build reports the canonical cwd block.
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface="unresolved",
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason="cwd_resolution_failed",
|
||||
path_snapshot=path_snapshot,
|
||||
)
|
||||
return original, trace
|
||||
|
||||
fallback_reason = ""
|
||||
assert binding is not None
|
||||
covers, executor_kind, executor_error = _executor_covers_kind(ctx, work_dir)
|
||||
if covers and executor_kind != "local":
|
||||
# R2/Q2-3: only a non-local backend skips the ladder — bare ``node``
|
||||
# resolves inside the container and a host path must never leak there.
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface="executor",
|
||||
environment="backend_path",
|
||||
reason="executor_backend_node",
|
||||
path_snapshot=path_snapshot,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
if executor_error:
|
||||
fallback_reason = executor_error
|
||||
|
||||
surface = _surface_for(ctx, binding, constraint)
|
||||
probe_token = requested if trigger == "runtime" else "node"
|
||||
located = shutil.which(probe_token) or ""
|
||||
if located and not os.path.isabs(located):
|
||||
# A relative PATH entry resolves against the WORKER cwd here but against
|
||||
# the command's work_dir at exec time: neither health nor brokenness is
|
||||
# provable from this process, so never substitute on that evidence —
|
||||
# run as written (argv and child env stay byte-identical) (T10).
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="host_path",
|
||||
reason="path_node_relative_entry_unprovable",
|
||||
fallback_reason=fallback_reason,
|
||||
path_snapshot=path_snapshot,
|
||||
runtime_path=located,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
if located:
|
||||
path_health = node_runtime_health(located)
|
||||
if path_health.healthy:
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="host_path",
|
||||
reason="path_node_healthy",
|
||||
fallback_reason=fallback_reason,
|
||||
path_snapshot=path_snapshot,
|
||||
runtime_path=located,
|
||||
runtime_version=path_health.version,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
path_fact = f"path_node_broken:{path_health.reason}:{located}"
|
||||
else:
|
||||
path_fact = f"path_node_missing:{probe_token}"
|
||||
|
||||
bundled = resolve_bundled_node() or ""
|
||||
bundled_health = node_runtime_health(bundled) if bundled else None
|
||||
if bundled_health is not None and bundled_health.healthy:
|
||||
prepend_dir = str(pathlib.Path(bundled).parent)
|
||||
resolved = bundled if trigger == "runtime" else requested
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=resolved,
|
||||
surface=surface,
|
||||
environment="bundled_node",
|
||||
reason="bundled_node_fallback",
|
||||
fallback_reason=path_fact,
|
||||
path_snapshot=path_snapshot,
|
||||
env_path_prepend=prepend_dir,
|
||||
runtime_path=bundled,
|
||||
runtime_version=bundled_health.version,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
if trigger == "runtime" and name != "verify_and_record":
|
||||
return _replace_request(name, original, argv, bundled), trace
|
||||
# verify_and_record keeps its ORIGINAL check text (receipt identity,
|
||||
# R4); the handler executes the resolved argv from this attestation.
|
||||
return original, trace
|
||||
|
||||
if bundled_health is not None:
|
||||
bundled_fact = f"bundled_node_broken:{bundled_health.reason}"
|
||||
else:
|
||||
bundled_fact = "bundled_node_missing"
|
||||
trace = _node_trace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="target_path",
|
||||
reason="no_usable_node",
|
||||
fallback_reason=f"{path_fact};{bundled_fact}",
|
||||
path_snapshot=path_snapshot,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
|
||||
|
||||
def active_node_resolution(ctx: Any) -> Optional[InterpreterResolutionTrace]:
|
||||
"""The node-family attestation of the CURRENT handler call, if any.
|
||||
|
||||
The registry scopes ``_active_interpreter_resolution`` to exactly one
|
||||
handler invocation, so a non-``None`` return always describes this call.
|
||||
"""
|
||||
|
||||
resolution = getattr(ctx, "_active_interpreter_resolution", None)
|
||||
if isinstance(resolution, InterpreterResolutionTrace) and resolution.family == "node":
|
||||
return resolution
|
||||
return None
|
||||
|
||||
|
||||
def interpreter_path_overlay(
|
||||
trace: Optional[InterpreterResolutionTrace],
|
||||
) -> Dict[str, str] | None:
|
||||
"""The ``{"PATH": ...}`` overlay attested by an emergency bundled fallback.
|
||||
|
||||
``None`` on every healthy/no-op resolution — callers must then leave their
|
||||
child env byte-identical to today's behavior.
|
||||
"""
|
||||
|
||||
if trace is None or not getattr(trace, "env_path_prepend", ""):
|
||||
return None
|
||||
base = trace.path_snapshot or os.environ.get("PATH", "")
|
||||
return {"PATH": trace.env_path_prepend + ((PATH_SEP + base) if base else "")}
|
||||
|
||||
|
||||
def apply_env_path_prepend(
|
||||
env: Mapping[str, str] | None,
|
||||
trace: Optional[InterpreterResolutionTrace],
|
||||
) -> Dict[str, str] | None:
|
||||
"""Child env with the attested bundled-runtime dir prepended to PATH.
|
||||
|
||||
Without an attested prepend the input is returned UNCHANGED (``None`` stays
|
||||
``None`` → inherit), keeping the healthy path byte-identical. With one, a
|
||||
``None``/inherit env becomes an explicit ``os.environ`` copy and PATH is
|
||||
rebuilt from the resolver's frozen snapshot. Never mutates process-global
|
||||
``os.environ``.
|
||||
"""
|
||||
|
||||
overlay = interpreter_path_overlay(trace)
|
||||
if overlay is None:
|
||||
return env if env is None else dict(env)
|
||||
out = dict(os.environ if env is None else env)
|
||||
if IS_WINDOWS:
|
||||
for key in [k for k in out if k.upper() == "PATH" and k != "PATH"]:
|
||||
del out[key]
|
||||
out["PATH"] = overlay["PATH"]
|
||||
return out
|
||||
|
||||
|
||||
_RESOLUTION_EVENT_TYPES = {
|
||||
"python": "python_interpreter_resolution",
|
||||
"node": "node_runtime_resolution",
|
||||
}
|
||||
|
||||
|
||||
def resolve_node_postgates(
|
||||
ctx: Any,
|
||||
tool_name: str,
|
||||
args: "Dict[str, Any]",
|
||||
*,
|
||||
runtime_mode: str,
|
||||
effective_constraint: Any = None,
|
||||
resolved_binding: Any = None,
|
||||
) -> "tuple[Dict[str, Any], Any]":
|
||||
"""Registry seam: resolve node once post-gates and record the trace.
|
||||
|
||||
Post-gates on purpose — the health check EXECUTES the candidate; the full
|
||||
placement rationale lives on ``resolve_process_node``.
|
||||
"""
|
||||
args, node_resolution = resolve_process_node(
|
||||
ctx,
|
||||
tool_name,
|
||||
args,
|
||||
runtime_mode=runtime_mode,
|
||||
effective_constraint=effective_constraint,
|
||||
resolved_binding=resolved_binding,
|
||||
)
|
||||
record_interpreter_resolution(ctx, node_resolution)
|
||||
return args, node_resolution
|
||||
|
||||
|
||||
@contextlib.contextmanager
|
||||
def interpreter_attestation(ctx: Any, trace: Any):
|
||||
"""Scope the per-call interpreter attestation to one handler invocation.
|
||||
|
||||
Publishes BOTH slots the downstream consumers read — the trace itself
|
||||
(``ctx._active_interpreter_resolution``: run_script's allowlist attestation,
|
||||
verify_and_record's R4 substitution, the handlers' emergency env prepend)
|
||||
and, exactly when the resolver substituted execution (an argv rewrite or an
|
||||
attested emergency PATH prepend), the ONE string slot
|
||||
``ctx._process_resolved_runtime`` that result_meta and the verify receipt
|
||||
disclose. Both slots are restored on exit whatever the handler did; on the
|
||||
healthy path the runtime slot is never set at all.
|
||||
"""
|
||||
missing = object()
|
||||
prior = getattr(ctx, "_active_interpreter_resolution", missing)
|
||||
ctx._active_interpreter_resolution = trace
|
||||
prior_runtime = getattr(ctx, "_process_resolved_runtime", missing)
|
||||
resolved_runtime = ""
|
||||
if trace is not None and (
|
||||
getattr(trace, "changed", False) or getattr(trace, "env_path_prepend", "")
|
||||
):
|
||||
resolved_runtime = str(
|
||||
getattr(trace, "runtime_path", "")
|
||||
or getattr(trace, "resolved_interpreter", "")
|
||||
or ""
|
||||
)
|
||||
if resolved_runtime:
|
||||
ctx._process_resolved_runtime = resolved_runtime
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
if prior is missing:
|
||||
try:
|
||||
delattr(ctx, "_active_interpreter_resolution")
|
||||
except AttributeError:
|
||||
pass
|
||||
else:
|
||||
ctx._active_interpreter_resolution = prior
|
||||
if resolved_runtime:
|
||||
if prior_runtime is missing:
|
||||
try:
|
||||
delattr(ctx, "_process_resolved_runtime")
|
||||
except AttributeError:
|
||||
pass
|
||||
else:
|
||||
ctx._process_resolved_runtime = prior_runtime
|
||||
|
||||
|
||||
def record_interpreter_resolution(ctx: Any, trace: Optional[InterpreterResolutionTrace]) -> None:
|
||||
"""Persist a compact, secret-free trace in the existing events log."""
|
||||
|
||||
if trace is None:
|
||||
return
|
||||
try:
|
||||
event: Dict[str, Any] = {
|
||||
"ts": utc_now_iso(),
|
||||
"type": _RESOLUTION_EVENT_TYPES.get(trace.family, "interpreter_resolution"),
|
||||
"task_id": str(getattr(ctx, "task_id", "") or ""),
|
||||
**trace.to_event(),
|
||||
}
|
||||
metadata = getattr(ctx, "task_metadata", {})
|
||||
if isinstance(metadata, dict):
|
||||
for key in ("root_task_id", "parent_task_id", "delegation_role"):
|
||||
value = metadata.get(key)
|
||||
if value not in (None, ""):
|
||||
event[key] = value
|
||||
correlation = getattr(ctx, "_current_llm_call_meta", {})
|
||||
if isinstance(correlation, dict):
|
||||
for key in ("execution_id", "round_id", "llm_call_id"):
|
||||
if correlation.get(key):
|
||||
event[key] = correlation[key]
|
||||
drive_logs = getattr(ctx, "drive_logs", None)
|
||||
if callable(drive_logs):
|
||||
log_dir = pathlib.Path(drive_logs())
|
||||
else:
|
||||
log_dir = pathlib.Path(getattr(ctx, "drive_root")) / "logs"
|
||||
append_jsonl(log_dir / "events.jsonl", event)
|
||||
except Exception:
|
||||
# Trace persistence must not make an otherwise-valid process call fail.
|
||||
return
|
||||
|
||||
|
||||
# Historic recorder name, kept for existing callers; the event type is chosen
|
||||
# by the trace's family either way.
|
||||
record_python_resolution = record_interpreter_resolution
|
||||
|
||||
|
||||
__all__ = [
|
||||
"InterpreterResolutionTrace",
|
||||
"PythonResolutionTrace",
|
||||
"active_node_resolution",
|
||||
"apply_env_path_prepend",
|
||||
"interpreter_attestation",
|
||||
"resolve_node_postgates",
|
||||
"interpreter_path_overlay",
|
||||
"record_interpreter_resolution",
|
||||
"record_python_resolution",
|
||||
"resolve_process_node",
|
||||
"resolve_process_python",
|
||||
]
|
||||
|
|
@ -19,7 +19,7 @@ from typing import Any, Dict, Iterable, List, Optional
|
|||
|
||||
from ouroboros.platform_layer import acquire_exclusive_file_lock, release_exclusive_file_lock
|
||||
from ouroboros.task_finalization import TERMINAL_ORIGIN_HOST_SALVAGE
|
||||
from ouroboros.utils import append_jsonl, iter_jsonl_objects, jsonl_append_lock_path, replace_atomic, utc_now_iso
|
||||
from ouroboros.utils import append_jsonl, iter_jsonl_objects, jsonl_append_lock_path, replace_atomic, strip_markdown, utc_now_iso
|
||||
|
||||
_ANNOTATIONS_NAME = "chat_annotations.jsonl"
|
||||
_COMPACT_AT_BYTES = 800_000
|
||||
|
|
@ -290,8 +290,23 @@ def append_chat_annotation(
|
|||
detail: str = "",
|
||||
options: Any = None,
|
||||
attachment_manifest: Any = None,
|
||||
require_latest_status: Any = None,
|
||||
require_latest_token: Any = None,
|
||||
) -> bool:
|
||||
"""Append one compact UI annotation; no semantic routing state is stored."""
|
||||
"""Append one compact UI annotation.
|
||||
|
||||
Presentation-first with ONE named exception (#198): a routing refusal row
|
||||
(status=needs_manual_target) is also the picker's durable decision-card
|
||||
authority — its token+options validate the owner's click, and the
|
||||
dispatch_pending/closing rows carry the click's first-wins/idempotency
|
||||
facts. Routing STATE still lives in the supervisor receipts (task-result
|
||||
admission, mailbox); the sidecar only arbitrates the card.
|
||||
|
||||
``require_latest_status`` (a set of status strings) turns the append into
|
||||
a compare-and-append under the annotations lock: the row is written only
|
||||
while the message's CURRENT latest status is in the set — the first-wins
|
||||
claim seam of the routing picker (#198). Absent/None keeps plain append.
|
||||
"""
|
||||
message_id = str(client_message_id or "").strip()
|
||||
if not message_id:
|
||||
return False
|
||||
|
|
@ -324,6 +339,15 @@ def append_chat_annotation(
|
|||
if lock_fd is None:
|
||||
return False
|
||||
try:
|
||||
if require_latest_status is not None or require_latest_token is not None:
|
||||
latest = _latest_annotations(path).get(message_id)
|
||||
if latest is not None:
|
||||
latest_status = str(latest.get("status") or "")
|
||||
latest_token = str(latest.get("routing_token") or "")
|
||||
if require_latest_status is not None and latest_status not in set(require_latest_status):
|
||||
return False # lost the claim race — the caller reads the truth back
|
||||
if require_latest_token is not None and latest_token not in set(require_latest_token):
|
||||
return False # a NEWER routing attempt owns the card now
|
||||
data = (json.dumps(row, ensure_ascii=False) + "\n").encode("utf-8")
|
||||
fd = os.open(str(path), os.O_WRONLY | os.O_CREAT | os.O_APPEND, 0o644)
|
||||
try:
|
||||
|
|
@ -379,6 +403,21 @@ def routing_options_with_labels(drive_root: Any, options: Any) -> List[Dict[str,
|
|||
return rows
|
||||
|
||||
|
||||
def routing_option_label(option: Any) -> str:
|
||||
"""One human label per manual-routing option — the HOST SSOT (the durable
|
||||
routing_options history row and the Telegram skill both render through it;
|
||||
web mirrors it as chat_activity.routingOptionLabel)."""
|
||||
if not isinstance(option, dict):
|
||||
return ""
|
||||
if str(option.get("label") or "").strip():
|
||||
return str(option["label"]).strip()
|
||||
if str(option.get("action") or "") == "new_task_in_project":
|
||||
return f"New task in {str(option.get('project_name') or 'Project')}"
|
||||
if option.get("title") or option.get("project_name"):
|
||||
return str(option.get("title") or option.get("project_name"))
|
||||
return "Project" if option.get("project_id") and not option.get("task_id") else "Task"
|
||||
|
||||
|
||||
def completion_status_label(result: Dict[str, Any], event: Dict[str, Any]) -> str:
|
||||
from ouroboros.task_results import (
|
||||
STATUS_CANCELLED, STATUS_COMPLETED, STATUS_FAILED, STATUS_REJECTED_DUPLICATE,
|
||||
|
|
@ -731,10 +770,17 @@ def append_terminal_task_projection(
|
|||
|
||||
|
||||
def _completion_excerpt(result: Dict[str, Any]) -> str:
|
||||
"""One plain-text excerpt for BOTH lifecycle writers (event + task_summary).
|
||||
|
||||
Markdown markers are stripped BEFORE whitespace flattening: the stripper's
|
||||
line-anchored heading/list patterns need the original newlines, and a
|
||||
flatten-first order would glue a ``##`` mid-line where no pattern (and no
|
||||
renderer) can treat it as markup again.
|
||||
"""
|
||||
if str(result.get("terminal_origin") or "") == TERMINAL_ORIGIN_HOST_SALVAGE:
|
||||
return ""
|
||||
for key in ("summary", "result", "error"):
|
||||
text = " ".join(str(result.get(key) or "").split())
|
||||
text = " ".join(strip_markdown(str(result.get(key) or "")).split())
|
||||
if text:
|
||||
return text if len(text) <= 240 else text[:239].rstrip() + "…"
|
||||
return ""
|
||||
|
|
|
|||
|
|
@ -1,422 +0,0 @@
|
|||
"""Surface-aware Python selection for host process tools.
|
||||
|
||||
Only the four public process launch surfaces opt into this resolver. It runs
|
||||
once at registry pre-dispatch so the deterministic guards and the handler see
|
||||
the same argv; launchers must not rewrite the interpreter afterwards.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import shutil
|
||||
import sys
|
||||
from dataclasses import asdict, dataclass
|
||||
from typing import Any, Dict, Mapping, Optional
|
||||
|
||||
from ouroboros.contracts.task_constraint import (
|
||||
TaskConstraint,
|
||||
normalize_task_constraint,
|
||||
)
|
||||
from ouroboros.platform_layer import project_venv_python
|
||||
from ouroboros.shell_parse import normalize_check_argv
|
||||
from ouroboros.tool_access import (
|
||||
ResolvedResourceBinding,
|
||||
build_resolved_resource_binding,
|
||||
path_is_relative_to,
|
||||
)
|
||||
from ouroboros.utils import append_jsonl, utc_now_iso
|
||||
|
||||
_PYTHON_TOKENS = frozenset({"python", "python3"})
|
||||
_PROCESS_TOOLS = frozenset({"run_command", "run_script", "start_service", "verify_and_record"})
|
||||
# Mirrors the run-kind set of tools/verify.py (`_RUN_KINDS`); keep in sync when
|
||||
# a new verify run-kind is introduced there.
|
||||
_VERIFY_RUN_KINDS = frozenset({"visible_verifier", "explicit_command", "explicit_metric"})
|
||||
_VERIFIED_RESOLUTIONS = frozenset(
|
||||
{
|
||||
("reviewed_skill_environment", "isolated_skill"),
|
||||
("executor_backend_python3", "backend_path"),
|
||||
("project_venv", "project_venv"),
|
||||
("agent_python", "ouroboros_agent"),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PythonResolutionTrace:
|
||||
tool: str
|
||||
requested_interpreter: str
|
||||
resolved_interpreter: str
|
||||
surface: str
|
||||
environment: str
|
||||
reason: str
|
||||
fallback_reason: str = ""
|
||||
error_reason: str = ""
|
||||
target_root: str = ""
|
||||
target_cwd: str = ""
|
||||
target_source: str = ""
|
||||
target_skill: str = ""
|
||||
|
||||
@property
|
||||
def changed(self) -> bool:
|
||||
return self.requested_interpreter != self.resolved_interpreter
|
||||
|
||||
@property
|
||||
def verified(self) -> bool:
|
||||
"""Whether the resolver proved the selected interpreter provenance."""
|
||||
|
||||
return (self.reason, self.environment) in _VERIFIED_RESOLUTIONS
|
||||
|
||||
def to_event(self) -> Dict[str, Any]:
|
||||
return {**asdict(self), "changed": self.changed}
|
||||
|
||||
|
||||
def _python_request(tool_name: str, args: Mapping[str, Any]) -> tuple[str, list[str] | None]:
|
||||
"""Return the exact eligible token and normalized argv, when applicable."""
|
||||
|
||||
if tool_name in {"run_command", "start_service"}:
|
||||
raw = args.get("cmd")
|
||||
if not isinstance(raw, list) or not raw:
|
||||
return "", None
|
||||
argv = [str(part) for part in raw]
|
||||
requested = str(argv[0]).strip()
|
||||
return (requested, argv) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
if tool_name == "run_script":
|
||||
requested = str(args.get("interpreter") or "python3").strip() or "python3"
|
||||
return (requested, None) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
if tool_name == "verify_and_record":
|
||||
kind = str(args.get("contract_kind") or "").strip()
|
||||
if kind not in _VERIFY_RUN_KINDS:
|
||||
return "", None
|
||||
argv = normalize_check_argv(args.get("check")) or []
|
||||
if not argv:
|
||||
return "", None
|
||||
requested = str(argv[0]).strip()
|
||||
return (requested, argv) if requested in _PYTHON_TOKENS else ("", None)
|
||||
|
||||
return "", None
|
||||
|
||||
|
||||
def _usable_executable(path_text: str) -> str:
|
||||
"""Validate an interpreter while preserving venv symlink semantics."""
|
||||
|
||||
text = str(path_text or "").strip()
|
||||
if not text:
|
||||
return ""
|
||||
candidate = pathlib.Path(text).expanduser()
|
||||
if not candidate.is_absolute():
|
||||
located = shutil.which(text)
|
||||
if not located:
|
||||
return ""
|
||||
candidate = pathlib.Path(located)
|
||||
try:
|
||||
if not candidate.is_file() or not os.access(candidate, os.X_OK):
|
||||
return ""
|
||||
except OSError:
|
||||
return ""
|
||||
# Do not resolve a venv's python symlink: executing the lexical path is what
|
||||
# lets Python discover the adjacent pyvenv.cfg and preserve the environment.
|
||||
return os.path.abspath(os.fspath(candidate))
|
||||
|
||||
|
||||
def _reviewed_skill_python(
|
||||
ctx: Any,
|
||||
binding: ResolvedResourceBinding | None = None,
|
||||
) -> tuple[str, str]:
|
||||
"""Return the lifecycle-proven isolated Python for the selected skill.
|
||||
|
||||
A dispatch binding is authoritative and loads exactly its physical payload
|
||||
against its canonical state root. Legacy task metadata is consulted only
|
||||
when a non-registry/direct caller supplied no binding.
|
||||
"""
|
||||
|
||||
try:
|
||||
from ouroboros.marketplace.isolated_deps import python_runtime_binary, read_deps_state
|
||||
from ouroboros.skill_loader import find_skill, load_skill
|
||||
from ouroboros.skill_readiness import skill_readiness_for_execution
|
||||
|
||||
if binding is not None:
|
||||
if binding.root != "skill_payload":
|
||||
return "", ""
|
||||
drive_root = pathlib.Path(binding.state_drive_root)
|
||||
loaded = load_skill(pathlib.Path(binding.base_path), drive_root)
|
||||
if loaded is None or loaded.name != binding.skill_name:
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
else:
|
||||
metadata = getattr(ctx, "task_metadata", {})
|
||||
metadata = metadata if isinstance(metadata, dict) else {}
|
||||
skill_name = str(metadata.get("skill") or "").strip()
|
||||
if not skill_name:
|
||||
return "", ""
|
||||
drive_root = pathlib.Path(getattr(ctx, "drive_root"))
|
||||
loaded = find_skill(drive_root, skill_name)
|
||||
if loaded is None or not skill_readiness_for_execution(drive_root, loaded).ready:
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
deps_state = read_deps_state(drive_root, loaded.name, loaded.skill_dir)
|
||||
if str(deps_state.get("status") or "") != "installed":
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
candidate = python_runtime_binary(loaded.skill_dir)
|
||||
usable = _usable_executable(str(candidate or ""))
|
||||
if usable:
|
||||
return usable, ""
|
||||
except Exception:
|
||||
return "", "reviewed_skill_environment_probe_failed"
|
||||
return "", "reviewed_skill_environment_unavailable"
|
||||
|
||||
|
||||
def _executor_covers(ctx: Any, work_dir: pathlib.Path) -> tuple[bool, str]:
|
||||
try:
|
||||
from ouroboros.workspace_executor import executor_ref_from_ctx, map_host_path
|
||||
|
||||
executor = executor_ref_from_ctx(ctx)
|
||||
if executor is None:
|
||||
return False, ""
|
||||
map_host_path(executor, pathlib.Path(work_dir).resolve(strict=False))
|
||||
return True, ""
|
||||
except ValueError:
|
||||
return False, ""
|
||||
except Exception:
|
||||
return False, "executor_resolution_failed"
|
||||
|
||||
|
||||
def _surface_for(
|
||||
ctx: Any,
|
||||
binding: ResolvedResourceBinding,
|
||||
constraint: Optional[TaskConstraint],
|
||||
) -> str:
|
||||
if binding.root != "active_workspace":
|
||||
return binding.root or "unresolved"
|
||||
if constraint and constraint.mode == "acting_subagent" and constraint.surface == "self_worktree":
|
||||
return "system_repo"
|
||||
mode = str(getattr(ctx, "workspace_mode", "") or "").strip().lower()
|
||||
if mode in {"external", "external_workspace", "genesis"}:
|
||||
return "external_workspace"
|
||||
system_repo = pathlib.Path(
|
||||
getattr(ctx, "system_repo_dir", None) or getattr(ctx, "repo_dir", binding.target_path)
|
||||
).resolve(strict=False)
|
||||
return "system_repo" if path_is_relative_to(binding.target_path, system_repo) else "external_workspace"
|
||||
|
||||
|
||||
def _trace_target(binding: ResolvedResourceBinding) -> Dict[str, str]:
|
||||
return {
|
||||
"target_root": binding.root,
|
||||
"target_cwd": str(binding.target_path),
|
||||
"target_source": binding.source,
|
||||
"target_skill": binding.skill_name,
|
||||
}
|
||||
|
||||
|
||||
def _project_root(ctx: Any, surface: str, work_dir: pathlib.Path) -> pathlib.Path:
|
||||
if surface == "external_workspace":
|
||||
workspace_root = getattr(ctx, "workspace_root", None)
|
||||
if workspace_root:
|
||||
candidate = pathlib.Path(workspace_root).resolve(strict=False)
|
||||
if path_is_relative_to(work_dir, candidate):
|
||||
return candidate
|
||||
return work_dir
|
||||
|
||||
|
||||
def _replace_request(
|
||||
tool_name: str,
|
||||
args: Mapping[str, Any],
|
||||
argv: list[str] | None,
|
||||
resolved: str,
|
||||
) -> Dict[str, Any]:
|
||||
out = dict(args)
|
||||
if tool_name in {"run_command", "start_service"}:
|
||||
new_argv = list(argv or [])
|
||||
new_argv[0] = resolved
|
||||
out["cmd"] = new_argv
|
||||
elif tool_name == "run_script":
|
||||
out["interpreter"] = resolved
|
||||
elif tool_name == "verify_and_record":
|
||||
new_argv = list(argv or [])
|
||||
new_argv[0] = resolved
|
||||
out["check"] = new_argv
|
||||
return out
|
||||
|
||||
|
||||
def resolve_process_python(
|
||||
ctx: Any,
|
||||
tool_name: str,
|
||||
args: Mapping[str, Any],
|
||||
*,
|
||||
runtime_mode: str,
|
||||
effective_constraint: Optional[TaskConstraint] = None,
|
||||
resolved_binding: ResolvedResourceBinding | None = None,
|
||||
) -> tuple[Dict[str, Any], Optional[PythonResolutionTrace]]:
|
||||
"""Resolve an exact ``python``/``python3`` request for one process tool."""
|
||||
|
||||
name = str(tool_name or "").strip()
|
||||
original = dict(args or {})
|
||||
if name not in _PROCESS_TOOLS:
|
||||
return original, None
|
||||
requested, argv = _python_request(name, original)
|
||||
if not requested:
|
||||
return original, None
|
||||
|
||||
constraint = normalize_task_constraint(effective_constraint)
|
||||
cwd_text = str(original.get("cwd") or "")
|
||||
binding = resolved_binding
|
||||
try:
|
||||
operation = "service" if name == "start_service" else "shell"
|
||||
if binding is None:
|
||||
binding = build_resolved_resource_binding(
|
||||
ctx,
|
||||
operation=operation,
|
||||
process_cwd=cwd_text,
|
||||
bucket=str(original.get("bucket") or ""),
|
||||
skill_name=str(original.get("skill_name") or ""),
|
||||
)
|
||||
work_dir = pathlib.Path(binding.target_path).resolve(strict=False)
|
||||
except Exception:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface="unresolved",
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason="cwd_resolution_failed",
|
||||
error_reason="cwd_resolution_failed",
|
||||
)
|
||||
return original, trace
|
||||
|
||||
fallback_reason = ""
|
||||
assert binding is not None
|
||||
skill_binding = resolved_binding
|
||||
if skill_binding is None and binding.root == "skill_payload":
|
||||
skill_binding = binding
|
||||
skill_python, skill_reason = _reviewed_skill_python(ctx, skill_binding)
|
||||
if skill_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=skill_python,
|
||||
surface="reviewed_skill",
|
||||
environment="isolated_skill",
|
||||
reason="reviewed_skill_environment",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, skill_python), trace
|
||||
if skill_reason:
|
||||
fallback_reason = skill_reason
|
||||
|
||||
executor_active, executor_error = _executor_covers(ctx, work_dir)
|
||||
if executor_active:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter="python3",
|
||||
surface="executor",
|
||||
environment="backend_path",
|
||||
reason="executor_backend_python3",
|
||||
fallback_reason=fallback_reason,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, "python3"), trace
|
||||
if executor_error and not fallback_reason:
|
||||
fallback_reason = executor_error
|
||||
|
||||
surface = _surface_for(ctx, binding, constraint)
|
||||
if surface in {"external_workspace", "user_files"}:
|
||||
project_python = project_venv_python(_project_root(ctx, surface, work_dir))
|
||||
if project_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=project_python,
|
||||
surface=surface,
|
||||
environment="project_venv",
|
||||
reason="project_venv",
|
||||
fallback_reason=fallback_reason,
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, project_python), trace
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason=fallback_reason or "project_venv_unavailable",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
|
||||
configured_agent_python = _usable_executable(
|
||||
os.environ.get("OUROBOROS_AGENT_PYTHON", "")
|
||||
)
|
||||
agent_python = configured_agent_python or _usable_executable(sys.executable or "")
|
||||
if agent_python:
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=agent_python,
|
||||
surface=surface,
|
||||
environment="ouroboros_agent",
|
||||
reason="agent_python",
|
||||
fallback_reason=(
|
||||
fallback_reason
|
||||
or ("agent_env_unavailable_process_fallback" if not configured_agent_python else "")
|
||||
),
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return _replace_request(name, original, argv, agent_python), trace
|
||||
|
||||
trace = PythonResolutionTrace(
|
||||
tool=name,
|
||||
requested_interpreter=requested,
|
||||
resolved_interpreter=requested,
|
||||
surface=surface,
|
||||
environment="target_path",
|
||||
reason="target_path_fallback",
|
||||
fallback_reason=fallback_reason or "agent_python_unavailable",
|
||||
error_reason="agent_python_unavailable",
|
||||
**_trace_target(binding),
|
||||
)
|
||||
return original, trace
|
||||
|
||||
|
||||
def record_python_resolution(ctx: Any, trace: Optional[PythonResolutionTrace]) -> None:
|
||||
"""Persist a compact, secret-free trace in the existing events log."""
|
||||
|
||||
if trace is None:
|
||||
return
|
||||
try:
|
||||
event: Dict[str, Any] = {
|
||||
"ts": utc_now_iso(),
|
||||
"type": "python_interpreter_resolution",
|
||||
"task_id": str(getattr(ctx, "task_id", "") or ""),
|
||||
**trace.to_event(),
|
||||
}
|
||||
metadata = getattr(ctx, "task_metadata", {})
|
||||
if isinstance(metadata, dict):
|
||||
for key in ("root_task_id", "parent_task_id", "delegation_role"):
|
||||
value = metadata.get(key)
|
||||
if value not in (None, ""):
|
||||
event[key] = value
|
||||
correlation = getattr(ctx, "_current_llm_call_meta", {})
|
||||
if isinstance(correlation, dict):
|
||||
for key in ("execution_id", "round_id", "llm_call_id"):
|
||||
if correlation.get(key):
|
||||
event[key] = correlation[key]
|
||||
drive_logs = getattr(ctx, "drive_logs", None)
|
||||
if callable(drive_logs):
|
||||
log_dir = pathlib.Path(drive_logs())
|
||||
else:
|
||||
log_dir = pathlib.Path(getattr(ctx, "drive_root")) / "logs"
|
||||
append_jsonl(log_dir / "events.jsonl", event)
|
||||
except Exception:
|
||||
# Trace persistence must not make an otherwise-valid process call fail.
|
||||
return
|
||||
|
||||
|
||||
__all__ = [
|
||||
"PythonResolutionTrace",
|
||||
"record_python_resolution",
|
||||
"resolve_process_python",
|
||||
]
|
||||
|
|
@ -679,9 +679,11 @@ def _classify_action(
|
|||
"reason_code": "provider_required_reasoning",
|
||||
})
|
||||
if effort_implicated and named_effort_value:
|
||||
from ouroboros.config import effort_one_step_down, effort_rank
|
||||
|
||||
next_effort = effort_one_step_down(current_effort)
|
||||
from ouroboros.config import EFFORT_SCALE, effort_one_step_down, effort_rank
|
||||
# QUOTED tiers inside [low, current) prescribe — even a negatively-quoted one (accepted FP); prose walks one rung.
|
||||
prescribed = [t for t in EFFORT_SCALE[effort_rank("low"):max(effort_rank(current_effort), 0)]
|
||||
if f"'{t}'" in low or f'"{t}"' in low]
|
||||
next_effort = prescribed[-1] if prescribed else effort_one_step_down(current_effort)
|
||||
if effort_rank(next_effort) >= effort_rank("low") and next_effort != current_effort:
|
||||
value_path = {
|
||||
"reasoning_effort": "reasoning_effort",
|
||||
|
|
|
|||
|
|
@ -7,6 +7,55 @@ from typing import Any, Callable, Dict, List
|
|||
from ouroboros.triad_review import parse_review_findings
|
||||
|
||||
|
||||
def contract_valid_actors(result: Any) -> List[Dict[str, Any]]:
|
||||
"""Actors with a DELIBERATE, CONTRACT-VALID reviewer object: parsed dict,
|
||||
recognizable verdict, parse_status not "malformed" — so a contract-DEMOTED or
|
||||
garbage response never votes (commit triad #1).
|
||||
|
||||
Owner ratification 2026-08-30: the acceptance-dialogue reducer now counts
|
||||
votes over ``_contributing_actors`` (a slot whose verdict did not reach the
|
||||
aggregate must not steer the loop either), and applies THIS predicate as the
|
||||
validity gate on top — the two are not the same test, and a hand-built or
|
||||
legacy row carrying a PASS/FAIL signal beside a malformed parse must still be
|
||||
unable to vote. It moved here from ``review_substrate`` with that change: this
|
||||
module is where the reviewer-row contract is decided, and the substrate sits
|
||||
exactly on its module-size cap."""
|
||||
from dataclasses import asdict
|
||||
|
||||
out: List[Dict[str, Any]] = []
|
||||
for actor in (getattr(result, "actors", None) or []):
|
||||
row = actor if isinstance(actor, dict) else asdict(actor)
|
||||
parsed = row.get("parsed")
|
||||
if str(row.get("parse_status") or "") == "malformed":
|
||||
continue
|
||||
if isinstance(parsed, dict) and str(
|
||||
parsed.get("verdict") or parsed.get("status") or ""
|
||||
).strip().upper() in {"PASS", "FAIL", "DEGRADED"}:
|
||||
out.append(row)
|
||||
return out
|
||||
|
||||
|
||||
def continue_vote_is_well_formed(parsed: Dict[str, Any]) -> bool:
|
||||
"""Whether a ``continue_actionable`` dialogue vote came with MATERIAL (owner
|
||||
ratification 2026-08-30, Rule 1; consumed by
|
||||
``review_substrate.aggregate_dialogue_status``).
|
||||
|
||||
Majority voting was rejected: one strong reviewer may still hold the
|
||||
acceptance loop open — but only while it can say what to do next. The same
|
||||
response must carry a concrete finding OR a completion_coach line; the
|
||||
correction-rail contract below already accepts coach-without-findings, so
|
||||
both count here. A bare "keep going" with neither is not a judgement the
|
||||
agent can act on: it bought one paid panel per round and never converged."""
|
||||
findings = parsed.get("findings")
|
||||
if isinstance(findings, list) and any(
|
||||
isinstance(item, dict)
|
||||
and str(item.get("item") or item.get("recommendation") or "").strip()
|
||||
for item in findings
|
||||
):
|
||||
return True
|
||||
return bool(str(parsed.get("completion_coach") or "").strip())
|
||||
|
||||
|
||||
def aggregate_review_actors(
|
||||
*,
|
||||
request: Any,
|
||||
|
|
|
|||
|
|
@ -144,6 +144,14 @@ def review_max_cycles() -> Optional[int]:
|
|||
return default_review_max_cycles()
|
||||
|
||||
|
||||
def review_max_cycles_source() -> str:
|
||||
"""Where the effective cap came from: ``owner_setting`` when the key is
|
||||
present in the environment (``config.apply_settings_to_env`` projects saved
|
||||
settings there, so an owner edit and an env override are the same fact),
|
||||
else ``shipped_default``. Provenance only — never a second parse."""
|
||||
return "owner_setting" if os.environ.get(REVIEW_MAX_CYCLES_KEY, "") else "shipped_default"
|
||||
|
||||
|
||||
def acceptance_max_improvement_passes_from_cycles() -> Optional[int]:
|
||||
"""Pure formula: task-acceptance improvement passes = shared cycles - 1
|
||||
(2 cycles → 1 pass); ``None`` when the shared cap is unlimited."""
|
||||
|
|
|
|||
|
|
@ -99,6 +99,9 @@ def task_acceptance_preclaim_refusal(ctx: Any) -> Any:
|
|||
ctx.tools._ctx,
|
||||
binding_hash=str((ctx.review_binding or {}).get("binding_hash") or ""),
|
||||
task_id=str(ctx.task_id or ""),
|
||||
# A-material: refuse a PAID dispatch whose material the tree already
|
||||
# bought, even when the binding hash moved (a cosmetic tool call moves it).
|
||||
paid_identity=str((ctx.review_binding or {}).get("paid_identity") or ""),
|
||||
)
|
||||
if projection.get("state") == "available" and not projection.get("binding_seen"):
|
||||
return None
|
||||
|
|
|
|||
|
|
@ -179,6 +179,7 @@ def build_task_acceptance_evidence(
|
|||
prov["mutation_attribution"] = "host_attested"
|
||||
from ouroboros.delegate_evidence import (
|
||||
acceptance_capability_deltas,
|
||||
acceptance_patch_dispositions,
|
||||
acceptance_substrate_facts,
|
||||
)
|
||||
|
||||
|
|
@ -188,6 +189,12 @@ def build_task_acceptance_evidence(
|
|||
if substrate_facts := acceptance_substrate_facts(ctx, task_id):
|
||||
ev["substrate_execution"] = redact_projection(substrate_facts).value
|
||||
prov["substrate_execution"] = "host_attested"
|
||||
# D-trace (owner 4=A): the parent's patch apply/reject attestations —
|
||||
# visibility for the panel, never a gate on apply. Absence = no
|
||||
# disposition recorded, not "reviewed clean".
|
||||
if patch_dispositions := acceptance_patch_dispositions(drive_root, task_id):
|
||||
ev["delegated_patch_dispositions"] = redact_projection(patch_dispositions).value
|
||||
prov["delegated_patch_dispositions"] = "host_attested"
|
||||
repo_diff = collect_turn_diff(ctx, include_recent_commit=include_recent_commit)
|
||||
diff_meta: Dict[str, Any] = {}
|
||||
if "OMISSION NOTE: truncated at " in str(repo_diff or "") or "... (truncated from " in str(repo_diff or ""):
|
||||
|
|
@ -871,6 +878,8 @@ from ouroboros.review_evidence_sections import ( # noqa: E402, F401 -- intentio
|
|||
_accept_claim_support_refs,
|
||||
_accept_effective_claims,
|
||||
_accept_enforce_budget,
|
||||
UNHASHED_ACCEPTANCE_DIALOGUE_HISTORY_KEY,
|
||||
UNHASHED_EVIDENCE_KEYS,
|
||||
_accept_obligation_row,
|
||||
_accept_owner_directives,
|
||||
_accept_protected_set,
|
||||
|
|
|
|||
|
|
@ -159,18 +159,40 @@ def _accept_obligation_row(o: Dict[str, Any]) -> Dict[str, Any]:
|
|||
row["previous_agent_reason"] = _accept_redact_cap(
|
||||
str(o.get("previous_reason")), 600,
|
||||
)
|
||||
# The LAST counter-argument in the exchange — without it this panel
|
||||
# cannot tell "already answered" from "never answered".
|
||||
if str(o.get("reviewer_rebuttal_response") or "").strip():
|
||||
row["previous_reviewer_response"] = _accept_redact_cap(
|
||||
str(o.get("reviewer_rebuttal_response")), 600,
|
||||
)
|
||||
return row
|
||||
|
||||
|
||||
# Reviewer-VISIBLE packet keys that are deliberately outside the packet's content
|
||||
# identity. The acceptance dialogue history is host-authored audit context that
|
||||
# grows by one row per panel: hashing it would shift the evidence revision — and
|
||||
# therefore mint a fresh paid binding — for a submission the agent did not change,
|
||||
# which is the acceptance pump A-material exists to close. Keep this set tiny; a
|
||||
# key belongs here only when it is derived from panels already paid for.
|
||||
UNHASHED_ACCEPTANCE_DIALOGUE_HISTORY_KEY = "acceptance_dialogue_history"
|
||||
UNHASHED_EVIDENCE_KEYS = (UNHASHED_ACCEPTANCE_DIALOGUE_HISTORY_KEY,)
|
||||
|
||||
|
||||
def task_acceptance_evidence_revision(evidence: Dict[str, Any]) -> str:
|
||||
"""Return the stable content revision used to bind acceptance evidence.
|
||||
|
||||
The evidence packet is already bounded and redacted by the shared builder.
|
||||
Hashing that exact packet lets the agent's cheap evidence call and the
|
||||
host-owned panel refer to the same revision without a second ledger.
|
||||
Hashing that exact packet — minus ``UNHASHED_EVIDENCE_KEYS`` — lets the
|
||||
agent's cheap evidence call and the host-owned panel refer to the same
|
||||
revision without a second ledger.
|
||||
"""
|
||||
packet = {
|
||||
key: value
|
||||
for key, value in (evidence or {}).items()
|
||||
if key not in UNHASHED_EVIDENCE_KEYS
|
||||
}
|
||||
payload = json.dumps(
|
||||
evidence or {},
|
||||
packet,
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
|
|
@ -324,6 +346,18 @@ def _accept_verification_summary(receipts: list) -> Dict[str, Any]:
|
|||
"latest_identity": _latest_identity,
|
||||
"latest_check": _latest_identity["check"],
|
||||
"latest_returncode": latest.get("returncode"),
|
||||
# Disclosure keys (node-runtime sprint, D6/R4), mirroring the fixed
|
||||
# ledger projection (`verification_receipt_ledger_row`) — a receipt key
|
||||
# missing from EITHER side is silently dropped, so both carry them.
|
||||
# Present only when the latest receipt has them: `latest_duration_ms`
|
||||
# (check process lifetime), `latest_signal` (POSIX signal name of a
|
||||
# killed check — a 9ms SIGKILL is not an ordinary red), and
|
||||
# `latest_resolved_runtime` (the substituted physical executable;
|
||||
# absent = the recorded check argv ran as written). The path is raw
|
||||
# host surface, so it goes through the redacting bound.
|
||||
**({"latest_duration_ms": latest.get("duration_ms")} if latest.get("duration_ms") is not None else {}),
|
||||
**({"latest_signal": str(latest.get("signal") or "")} if latest.get("signal") else {}),
|
||||
**({"latest_resolved_runtime": _accept_redact_cap(str(latest.get("resolved_runtime") or ""), 300)} if latest.get("resolved_runtime") else {}),
|
||||
"latest_expected_match": str(latest.get("expected_match") or ""),
|
||||
"latest_summary": _accept_redact_cap(str(latest.get("summary") or ""), 2000),
|
||||
# C: aggregate the after-only artifact-lifecycle flag across ALL receipts (a deleted
|
||||
|
|
|
|||
|
|
@ -262,10 +262,21 @@ class NativeToolRoundReviewExecutor(ReviewSlotExecutor):
|
|||
scratch = tempfile.mkdtemp(prefix="ouro-native-review-")
|
||||
registry = None
|
||||
total_usage: Dict[str, Any] = {}
|
||||
transcript_chars = len(self.episode_prompt)
|
||||
final_answer: Optional[str] = None
|
||||
try:
|
||||
registry, schemas = self._inspection_registry(root, scratch)
|
||||
# The counter measures what every send actually carries: the
|
||||
# system instructions and the tool schemas ride EVERY provider
|
||||
# call, and tool-call argument objects accumulate in `messages`
|
||||
# exactly like results do. Counting only prompt+content+results
|
||||
# understated each send by the fixed system/schema cost and let
|
||||
# the argument tail drift past the promised bound unmeasured.
|
||||
# Units are CHARS throughout — same as the cap.
|
||||
transcript_chars = (
|
||||
len(self.episode_prompt)
|
||||
+ len(_NATIVE_REVIEW_INSTRUCTIONS)
|
||||
+ len(json.dumps(schemas, ensure_ascii=False, default=str))
|
||||
)
|
||||
messages: List[Dict[str, Any]] = [
|
||||
{"role": "system", "content": _NATIVE_REVIEW_INSTRUCTIONS},
|
||||
{"role": "user", "content": self.episode_prompt},
|
||||
|
|
@ -337,6 +348,13 @@ class NativeToolRoundReviewExecutor(ReviewSlotExecutor):
|
|||
total_usage[_fact] = usage[_fact]
|
||||
content = str(msg.get("content") or "") if isinstance(msg, dict) else ""
|
||||
transcript_chars += len(content)
|
||||
# Tool-call objects (names + argument JSON) join `messages`
|
||||
# below and ride every later send — the cumulative half of
|
||||
# the previous under-count.
|
||||
transcript_chars += sum(
|
||||
len(json.dumps(tc, ensure_ascii=False, default=str))
|
||||
for tc in tool_calls if isinstance(tc, dict)
|
||||
)
|
||||
if content and not tool_calls:
|
||||
final_answer = content
|
||||
break
|
||||
|
|
@ -349,6 +367,8 @@ class NativeToolRoundReviewExecutor(ReviewSlotExecutor):
|
|||
assistant.setdefault("role", "assistant")
|
||||
messages.append(assistant)
|
||||
for tc in tool_calls:
|
||||
if not isinstance(tc, dict):
|
||||
continue # a non-dict tool_call is malformed provider output, not a crash
|
||||
call_id = str(tc.get("id") or "")
|
||||
function = tc.get("function") if isinstance(tc, dict) else {}
|
||||
name = str((function or {}).get("name") or "")
|
||||
|
|
|
|||
|
|
@ -784,12 +784,15 @@ from ouroboros.review_records import ( # noqa: E402, F401 -- intentional public
|
|||
|
||||
from ouroboros.review_verdict import ( # noqa: E402, F401 -- intentional public re-exports
|
||||
DIALOGUE_CONTINUE,
|
||||
DIALOGUE_INCONCLUSIVE,
|
||||
DIALOGUE_STABLE_DISAGREEMENT,
|
||||
DIALOGUE_STATUS_VALUES,
|
||||
DIALOGUE_TERMINAL_STATUSES,
|
||||
DIALOGUE_UNREACHABLE,
|
||||
DIALOGUE_VOTE_ABSTAIN_INVALID,
|
||||
DIALOGUE_VOTE_CONTINUE_WITHOUT_FINDINGS,
|
||||
_CRITERION_STATUSES,
|
||||
_TIER_ORDER,
|
||||
_contract_valid_actors,
|
||||
_contributing_actors,
|
||||
_criteria_have_supported_evidence,
|
||||
_criteria_shape_valid,
|
||||
|
|
|
|||
|
|
@ -142,54 +142,55 @@ DIALOGUE_STABLE_DISAGREEMENT = "stable_disagreement"
|
|||
DIALOGUE_STATUS_VALUES = (DIALOGUE_CONTINUE, DIALOGUE_UNREACHABLE, DIALOGUE_STABLE_DISAGREEMENT)
|
||||
|
||||
|
||||
def _contract_valid_actors(result: Any) -> List[Dict[str, Any]]:
|
||||
"""Actors with a DELIBERATE, CONTRACT-VALID reviewer object: parsed dict,
|
||||
recognizable verdict, parse_status not "malformed". Wider than
|
||||
``_contributing_actors`` (deliberate DEGRADED keeps its vote, sol #3) but a
|
||||
contract-DEMOTED/garbage response never votes terminal (commit triad #1)."""
|
||||
out: List[Dict[str, Any]] = []
|
||||
for actor in (getattr(result, "actors", None) or []):
|
||||
row = actor if isinstance(actor, dict) else asdict(actor)
|
||||
parsed = row.get("parsed")
|
||||
if str(row.get("parse_status") or "") == "malformed":
|
||||
continue
|
||||
if isinstance(parsed, dict) and str(
|
||||
parsed.get("verdict") or parsed.get("status") or ""
|
||||
).strip().upper() in {"PASS", "FAIL", "DEGRADED"}:
|
||||
out.append(row)
|
||||
return out
|
||||
DIALOGUE_TERMINAL_STATUSES = (DIALOGUE_UNREACHABLE, DIALOGUE_STABLE_DISAGREEMENT)
|
||||
# Reducer OUTPUTS, deliberately outside DIALOGUE_STATUS_VALUES (reviewer vocabulary
|
||||
# unchanged): no well-formed vote at all — neither terminal nor a licence to
|
||||
# continue — plus the buckets that keep a REFUSED vote disclosed, not dropped.
|
||||
DIALOGUE_INCONCLUSIVE = "inconclusive"
|
||||
DIALOGUE_VOTE_CONTINUE_WITHOUT_FINDINGS = "continue_without_findings"
|
||||
DIALOGUE_VOTE_ABSTAIN_INVALID = "abstain_invalid"
|
||||
|
||||
|
||||
def aggregate_dialogue_status(result: Any, *, quorum: int) -> Dict[str, Any]:
|
||||
"""Pure reducer over the reviewers' typed ``dialogue_status`` votes (A5, P5):
|
||||
the host validates the enum, applies the caller's quorum, and transports the
|
||||
result. Precedence: any continue vote from a QUORUM-CONTRIBUTING actor keeps
|
||||
the loop; else a quorum of terminal votes terminates. Missing/invalid votes
|
||||
default to ``continue_actionable`` (fail-safe, backward-compatible).
|
||||
Returns ``{"status", "votes"}`` with the full distribution for audit."""
|
||||
contributing = {str(a.get("slot_id", "")) for a in _contributing_actors(result)}
|
||||
"""Pure reducer over the reviewers' typed ``dialogue_status`` votes (A5, P5).
|
||||
Votes are counted over the CONTRIBUTING actors, gated by contract validity
|
||||
(owner ratification 2026-08-30, replacing the sol #3 widening). Precedence,
|
||||
over WELL-FORMED votes only: a continue keeps the loop; else >=1 terminal vote
|
||||
ends it with the unreachable-vs-disagreement tie-break; else ``inconclusive``,
|
||||
which grants the dialogue no authority either way. A continue WITHOUT material
|
||||
and a missing/invalid vote both ABSTAIN and stay disclosed in the distribution;
|
||||
neither may default the loop into another paid round — which is exactly what
|
||||
the old fail-safe ``continue`` default did. ``quorum`` no longer gates the
|
||||
terminal side (ending the dialogue is the cheap direction, and the two ends are
|
||||
now symmetric); it rides the record so it stays auditable beside the votes."""
|
||||
from ouroboros.review_actor_aggregation import contract_valid_actors, continue_vote_is_well_formed
|
||||
|
||||
valid_slots = {str(row.get("slot_id", "")) for row in contract_valid_actors(result)}
|
||||
votes: Dict[str, List[str]] = {}
|
||||
for row in _contract_valid_actors(result):
|
||||
for row in _contributing_actors(result):
|
||||
slot_id = str(row.get("slot_id", ""))
|
||||
if slot_id not in valid_slots:
|
||||
continue # a PASS/FAIL signal beside a malformed parse never votes
|
||||
parsed = row.get("parsed") if isinstance(row.get("parsed"), dict) else {}
|
||||
vote = str(parsed.get("dialogue_status") or "").strip().lower()
|
||||
if vote not in DIALOGUE_STATUS_VALUES:
|
||||
vote = DIALOGUE_CONTINUE
|
||||
votes.setdefault(vote, []).append(str(row.get("slot_id", "")))
|
||||
continue_slots = votes.get(DIALOGUE_CONTINUE, [])
|
||||
vote = DIALOGUE_VOTE_ABSTAIN_INVALID
|
||||
elif vote == DIALOGUE_CONTINUE and not continue_vote_is_well_formed(parsed):
|
||||
vote = DIALOGUE_VOTE_CONTINUE_WITHOUT_FINDINGS
|
||||
votes.setdefault(vote, []).append(slot_id)
|
||||
unreachable = votes.get(DIALOGUE_UNREACHABLE, [])
|
||||
disagreement = votes.get(DIALOGUE_STABLE_DISAGREEMENT, [])
|
||||
terminal = unreachable + disagreement
|
||||
if any(slot in contributing for slot in continue_slots):
|
||||
if votes.get(DIALOGUE_CONTINUE):
|
||||
status = DIALOGUE_CONTINUE
|
||||
elif len(terminal) >= max(1, int(quorum)):
|
||||
elif unreachable or disagreement:
|
||||
status = (
|
||||
DIALOGUE_UNREACHABLE
|
||||
if len(unreachable) >= len(disagreement)
|
||||
else DIALOGUE_STABLE_DISAGREEMENT
|
||||
)
|
||||
else:
|
||||
status = DIALOGUE_CONTINUE
|
||||
return {"status": status, "votes": votes}
|
||||
status = DIALOGUE_INCONCLUSIVE
|
||||
return {"status": status, "votes": votes, "quorum": max(1, int(quorum))}
|
||||
|
||||
|
||||
def _unresolved_evidence_ref_labels(run: Any) -> List[str]:
|
||||
|
|
|
|||
86
ouroboros/routing_wait.py
Normal file
86
ouroboros/routing_wait.py
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
"""Durable routing-receipt waits (SSOT for the tool layer AND the gateway).
|
||||
|
||||
Extracted verbatim from ``ouroboros/tools/control.py`` (#198): the routing
|
||||
picker's HTTP dispatcher must poll the SAME durable receipts the LLM routing
|
||||
tools poll — the task-result ``promotion_admission`` record and the exact
|
||||
chat-annotation receipt — without importing the whole tool registry into the
|
||||
gateway process. Root-parameterized; ``tools/control.py`` keeps thin wrappers
|
||||
that resolve the root from the tool context.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict
|
||||
|
||||
PROMOTE_CONFIRM_TIMEOUT_SEC = 15.0
|
||||
PROMOTE_CONFIRM_POLL_SEC = 0.05
|
||||
|
||||
|
||||
def wait_for_promotion_admission(
|
||||
root: Path,
|
||||
task_id: str,
|
||||
routing_token: str,
|
||||
*,
|
||||
client_message_id: str = "",
|
||||
timeout_sec: float = PROMOTE_CONFIRM_TIMEOUT_SEC,
|
||||
poll_sec: float = PROMOTE_CONFIRM_POLL_SEC,
|
||||
) -> Dict[str, Any]:
|
||||
"""Wait for matching-token admission in the canonical task-result SSOT."""
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
deadline = time.monotonic() + max(0.0, float(timeout_sec))
|
||||
while True:
|
||||
result = load_task_result(root, task_id) or {}
|
||||
admission = result.get("promotion_admission")
|
||||
if (
|
||||
isinstance(admission, dict)
|
||||
and str(admission.get("routing_token") or "") == routing_token
|
||||
):
|
||||
status = str(admission.get("status") or "")
|
||||
if status in {"scheduled", "rejected", "unconfirmed"}:
|
||||
return {**admission, "task_status": str(result.get("status") or "")}
|
||||
# A duplicate id must never overwrite the existing task_result merely
|
||||
# to report the loser. The exact-token chat annotation is therefore a
|
||||
# negative-only fallback; positive scheduling authority stays solely in
|
||||
# the task-result admission record.
|
||||
if str(client_message_id or "").strip():
|
||||
from ouroboros.project_dialogue import chat_annotation_receipt
|
||||
|
||||
receipt = chat_annotation_receipt(
|
||||
root, str(client_message_id), routing_token
|
||||
)
|
||||
if str(receipt.get("status") or "") in {
|
||||
"needs_manual_target",
|
||||
"rejected",
|
||||
"unconfirmed",
|
||||
}:
|
||||
return receipt
|
||||
if time.monotonic() >= deadline:
|
||||
return {"status": "unconfirmed", "reason": "confirmation_timeout"}
|
||||
time.sleep(poll_sec)
|
||||
|
||||
|
||||
def wait_for_routing_annotation(
|
||||
root: Path,
|
||||
client_message_id: str,
|
||||
routing_token: str,
|
||||
*,
|
||||
timeout_sec: float = PROMOTE_CONFIRM_TIMEOUT_SEC,
|
||||
poll_sec: float = PROMOTE_CONFIRM_POLL_SEC,
|
||||
) -> Dict[str, Any]:
|
||||
"""Wait for an exact existing chat-annotation receipt (manual/steer)."""
|
||||
from ouroboros.project_dialogue import chat_annotation_receipt
|
||||
|
||||
if not str(client_message_id or "").strip():
|
||||
return {"status": "unconfirmed", "reason": "client_message_id_missing"}
|
||||
deadline = time.monotonic() + max(0.0, float(timeout_sec))
|
||||
while True:
|
||||
receipt = chat_annotation_receipt(root, client_message_id, routing_token)
|
||||
status = str(receipt.get("status") or "")
|
||||
if status in {"delivered", "needs_manual_target", "unconfirmed"}:
|
||||
return receipt
|
||||
if time.monotonic() >= deadline:
|
||||
return {"status": "unconfirmed", "reason": "confirmation_timeout"}
|
||||
time.sleep(poll_sec)
|
||||
|
|
@ -134,10 +134,13 @@ TOOL_POLICY: Dict[str, str] = {
|
|||
"send_photo": POLICY_SKIP,
|
||||
"send_video": POLICY_SKIP,
|
||||
"send_file": POLICY_SKIP,
|
||||
# Structured links cannot exceed the task's existing owner-chat delivery authority.
|
||||
"send_links": POLICY_SKIP,
|
||||
"presence_finish": POLICY_SKIP,
|
||||
"presence_cancel_work": POLICY_SKIP,
|
||||
"configure_presence": POLICY_SKIP,
|
||||
"initiate_presence": POLICY_SKIP,
|
||||
"escalate": POLICY_SKIP,
|
||||
"forward_to_worker": POLICY_SKIP,
|
||||
"compact_context": POLICY_SKIP,
|
||||
"enable_tools": POLICY_SKIP,
|
||||
|
|
@ -278,7 +281,7 @@ def _normalize_resolved_python_subject(raw_cmd: Any, python_resolution: Any) ->
|
|||
literal ``python`` token for the existing module allowlist decision.
|
||||
"""
|
||||
|
||||
from ouroboros.python_interpreter import PythonResolutionTrace
|
||||
from ouroboros.process_interpreters import PythonResolutionTrace
|
||||
|
||||
if not isinstance(python_resolution, PythonResolutionTrace) or not python_resolution.verified:
|
||||
return ""
|
||||
|
|
|
|||
|
|
@ -89,6 +89,7 @@ def _periodic_supervisor_maintenance(last_custody_reap: list, last_review_reconc
|
|||
# its owning task is no longer running. It has no pid, so the process
|
||||
# reaper cannot see it — but it is still spending quota and still writing.
|
||||
_reconcile_delegated_runs(live_tasks)
|
||||
_cursor_refresh_settled_terminals()
|
||||
except Exception:
|
||||
log.debug("Periodic custody reap failed", exc_info=True)
|
||||
if time.time() - last_review_reconcile[0] > 300:
|
||||
|
|
@ -179,6 +180,25 @@ def _startup_prune_sweeps() -> None:
|
|||
log.debug("Headless task drive prune failed", exc_info=True)
|
||||
|
||||
|
||||
def _cursor_refresh_settled_terminals() -> None:
|
||||
"""Cursor-driven pass: runs settled OUTSIDE a generation's reconcile
|
||||
outcomes (terminal-boundary settlements, earlier generations) never
|
||||
reappear in the orphan sweep, so their tasks' stored evidence would stay
|
||||
stale forever. Bounded to newly appended custody rows per tick. At BOOT
|
||||
this runs AFTER the D1a backfill (see ``_startup_custody_sweep``), so a
|
||||
same-generation heal keeps its pinned ``boot_backfill`` attribution and
|
||||
the cursor's change-gated pass advances past it without a second write.
|
||||
"""
|
||||
try:
|
||||
from ouroboros.delegate_terminal import refresh_recently_settled_terminals
|
||||
|
||||
refreshed = refresh_recently_settled_terminals(DATA_DIR)
|
||||
if refreshed:
|
||||
log.info("Cursor refresh healed %d stale terminal result(s)", refreshed)
|
||||
except Exception:
|
||||
log.debug("Cursor terminal-refresh pass failed", exc_info=True)
|
||||
|
||||
|
||||
def _startup_custody_sweep() -> None:
|
||||
"""Both custody surfaces, swept once per generation at supervisor startup.
|
||||
|
||||
|
|
@ -194,6 +214,22 @@ def _startup_custody_sweep() -> None:
|
|||
except Exception:
|
||||
log.debug("Process custody startup reap failed", exc_info=True)
|
||||
_reconcile_delegated_runs(set())
|
||||
try:
|
||||
# D1a boot backfill, ONCE per generation and AFTER the orphan reconcile
|
||||
# (so this generation's settlements are already visible to the audit):
|
||||
# a run settled in a PREVIOUS generation never appears in any current
|
||||
# pass's outcomes, so the sweep-side refresh above can never reach its
|
||||
# task's stored disclosure — the backfill joins from the stored terminal
|
||||
# results instead and heals every generation-crossing stale row.
|
||||
from ouroboros.delegate_terminal import backfill_terminal_reconciliations
|
||||
|
||||
refreshed = backfill_terminal_reconciliations(DATA_DIR)
|
||||
if refreshed:
|
||||
log.info("Boot custody backfill refreshed %d stored disclosure(s): %s",
|
||||
len(refreshed), refreshed)
|
||||
except Exception:
|
||||
log.debug("Boot custody-disclosure backfill failed", exc_info=True)
|
||||
_cursor_refresh_settled_terminals()
|
||||
try:
|
||||
# Phase A boot migration: legacy ``cancel_requested`` status latches
|
||||
# become ordinary durable cancel intents; the supervisor watchdog then
|
||||
|
|
|
|||
|
|
@ -266,7 +266,7 @@ SETTINGS_DEFAULTS = {**UPDATE_SETTINGS_DEFAULTS,
|
|||
# writes at the documented 2x-vs-1.25x ratio). Non-Anthropic wire formats are a NO-OP by construction
|
||||
# (Gemini documents no ttl field — the v5.30.0 outage class).
|
||||
"OUROBOROS_PROMPT_CACHE_TTL": "1h",
|
||||
# Reasoning effort per task type: none | low | medium | high
|
||||
# Reasoning effort per task type: any EFFORT_SCALE tier (the ordered SSOT in settings_scales)
|
||||
"OUROBOROS_EFFORT_TASK": "medium",
|
||||
"OUROBOROS_EFFORT_EVOLUTION": "high",
|
||||
"OUROBOROS_EFFORT_REVIEW": "high",
|
||||
|
|
|
|||
|
|
@ -13,10 +13,10 @@ from typing import Any
|
|||
|
||||
from ouroboros.settings_defaults import SETTINGS_DEFAULTS
|
||||
|
||||
# v6.57.0 — EFFORT_SCALE: ORDERED reasoning-effort SSOT (low→high), the single place a tier
|
||||
# is defined (settings, llm.py builder, switch_model enum, subagent lanes). Exact-route
|
||||
# request-wire recovery, not legacy model-global evidence, owns provider adaptation.
|
||||
EFFORT_SCALE: tuple[str, ...] = ("none", "minimal", "low", "medium", "high", "xhigh", "max")
|
||||
# v6.57.0 — EFFORT_SCALE: ORDERED reasoning-effort SSOT (low→high), the single place a tier is
|
||||
# defined (settings, llm.py builder, switch_model enum, subagent lanes). `ultra` = the codex
|
||||
# vendor tier above `max`; above-ceiling tiers adapt per route (API wire recovery / delegated).
|
||||
EFFORT_SCALE: tuple[str, ...] = ("none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra")
|
||||
|
||||
|
||||
def effort_rank(value: str) -> int:
|
||||
|
|
|
|||
|
|
@ -119,10 +119,12 @@ BAND_PATHS = {
|
|||
"ouroboros/consciousness.py": "Durable Background Consciousness observation inbox and bounded truthful replay",
|
||||
"ouroboros/context.py": "Entered the band from the 1501-1600 zone (1590 lines) by the v7 D03 extraction of the runtime-section fact builders into ouroboros/context_runtime_facts.py; shrink-only residue of the split, not new growth.",
|
||||
"ouroboros/delegate_custody.py": "D07 DEL1 split brought the custody monolith DOWN from the 1600 hard cap into the band (1600->1305); reconcile family extracted to delegate_custody_reconcile.py, shrink-only direction",
|
||||
"ouroboros/extension_plugin_api.py": "F6 upstream sync: the T14 node-runtime companion rewrite (manifest-PATH override guard + bundled-node argv substitution) landed in the registration owner; shrink-only from here",
|
||||
"ouroboros/extension_process_runner.py": None,
|
||||
"ouroboros/gateway/control.py": "Entered the band from 966 lines: the update-flow redesign added the shared stash-first prologue (_stash_local_work_fenced/_unwind_stashed_update) and the review-wave affordability floor to the update apply orchestration (update-flow-redesign sprint, Q9/Q10 owner decisions).",
|
||||
"ouroboros/gateway/history.py": None,
|
||||
"ouroboros/gateway/settings.py": "Retiring persistent auto-Low removed the former giant debt; the remaining owner and reviewer settings endpoints stay centralized while tracked in the shrinking band.",
|
||||
"ouroboros/loop_acceptance_review.py": "F6 upstream sync: the A-material acceptance family (paid identity, free replay, identical-refusal terminal, dialogue history) folded into the campaign review leaf per the sync principle (upstream leaf acceptance_dialogue.py retired)",
|
||||
"ouroboros/loop_delivery.py": "F6 upstream sync: the delivery-protocol upstream deltas (hold-control literals, trailing-object/fence-aware protocol parsers) folded into the campaign delivery leaf (upstream leaf delivery_protocol.py retired)",
|
||||
"ouroboros/loop_forced_finalization.py": "Forced-finalization rail of the v7 L-B loop split: one cohesive owner for the forced/orphan/absorption path, moved byte-preserving from loop.py (D01 lane).",
|
||||
"ouroboros/loop_tool_execution.py": None,
|
||||
"ouroboros/marketplace/ouroboroshub.py": "Entered the band from 373 lines: the hubflow sprint added the adopt transaction (eligibility prelude, CAS re-verification, move-aside + state-quintet snapshot, verified rollback with per-step error collection, retention finalize) beside the existing install/update flows (hubflow sprint, adopt-in-ouroboroshub owner decision D4).",
|
||||
|
|
@ -141,7 +143,6 @@ BAND_PATHS = {
|
|||
"ouroboros/subagent_runtime.py": "Configured-retry refusals mirrored typed (triad 2026-08-30) push the module just over 1000; no new subsystem, same seam.",
|
||||
"ouroboros/subagent_worktrees.py": "Owner-sanctioned strict-registry delta (v7 rows 1083-1092, fork F-1=A) grew the module 1000->1082: typed refusal of a malformed registry instead of silent collapse-to-empty; shrink-only direction",
|
||||
"ouroboros/subagents.py": "D07 split brought the dispatch monolith DOWN from 1593 into the band (->1370); route-health family extracted to subagent_route_health.py, shrink-only direction",
|
||||
"ouroboros/task_results.py": "Authority reads need an explicit strict mode so malformed child records cannot become a false zero count.",
|
||||
"ouroboros/task_status.py": None,
|
||||
"ouroboros/tools/browser.py": None,
|
||||
"ouroboros/tools/claude_advisory_review.py": "F2.3b advisory re-derive: the 2279-line GIANT parent re-enters the band at 1434 after the preflight_review_prompt/preflight_review_run split (admission policy, native episode, size gates and the tool entries stay with the facade)",
|
||||
|
|
@ -154,7 +155,6 @@ BAND_PATHS = {
|
|||
"ouroboros/tools/review_context_atlas.py": "Grew INTO the band by the #284 pack-arithmetic fixes: measured render charged at admission, exact per-row costs, target capped at the hard rail, honest eviction diagnostics \u2014 all in the module that owns the arithmetic.",
|
||||
"ouroboros/tools/skill_exec.py": None,
|
||||
"ouroboros/tools/skill_publish.py": "Entered the band from 952 lines: publish now writes the OuroborosHub publication receipt at pr_opened through the shared locked-update seam and maps the receipt from the validated serialized form (hubflow sprint, receipt-as-only-stored-fact design).",
|
||||
"ouroboros/tools/subagent_integration.py": "D07 DEL1 split brought the integration monolith DOWN from 1599 into the band (->1027); delegated-disposition family extracted to subagent_integration_delegated.py, shrink-only direction",
|
||||
"ouroboros/usage_accounting.py": "Entered the band from the 1501-1600 zone (1600 lines) by the v7 L-C2 extraction of the one-time legacy usage import into ouroboros/usage_legacy_import.py; shrink-only residue of the split, not new growth.",
|
||||
"ouroboros/utils.py": None,
|
||||
"ouroboros/workspace_executor.py": None,
|
||||
|
|
@ -165,15 +165,14 @@ BAND_PATHS = {
|
|||
"supervisor/evolution_lifecycle.py": None,
|
||||
"supervisor/message_bus.py": "push_log gained the A2A frame suppression and explicit-addressing contract comments in the live log routing fix (fix/main-chat-leak sprint); the module sat at exactly 1000 lines before it",
|
||||
"supervisor/queue_transitions.py": "F2.2 cancel/custody organ: the owner-stop campaign closure (_close_campaign_after_owner_stop, reference row 970) moved in from the hot events monolith to live beside stop_evolution_tasks - one honesty rule, one owner; 1016 lines, shrink-only from here",
|
||||
"supervisor/task_reaper.py": "Entered the band from 907 lines: timeout-retry admission now serializes queue publication, reciprocal result lineage, cancellation-wins handoff, and failed-terminal-write custody in the existing off-loop reaper owner.",
|
||||
"supervisor/terminal_delivery.py": None,
|
||||
"supervisor/update_merge.py": "Entered the band from above (1593 lines) by extraction: the F2.4 update-engine re-split moved the planner, the clean-plan commit builder and the live materializer \u2014 the carrier engine's three insertion points \u2014 into supervisor/update_merge_plan.py (D34 return, owner answers 5.12-5.14=A); shrink-only.",
|
||||
"tests/test_acting_subagents.py": None,
|
||||
"tests/test_advisory_observability.py": None,
|
||||
"tests/test_build_scripts.py": None,
|
||||
"tests/test_commit_gate.py": None,
|
||||
"tests/test_contracts.py": None,
|
||||
"tests/test_cybergym_protocol.py": "CyberGym protocol suite arrived in one piece with the benchmark (drift heal); split when the next protocol family lands.",
|
||||
"tests/test_delegate_answer.py": "Entered the band by the #204 escalation-route pins (walk-up, schema and expiry-note source pins) on top of the phase-B interaction suite; one coherent delegated-question surface, split only when a natural seam appears",
|
||||
"tests/test_delegated_skill_payload.py": "Sol scope-review fix batch: P1 trust probes (forged index, symlinked git metadata), P2 golden-E2E review close and schema/docs pins joined the existing R1+gate-fix payload suite.",
|
||||
"tests/test_evolution_redesign.py": None,
|
||||
"tests/test_external_review_script.py": "F2.3b D31 port: the reference trusted-base pin suite (probe-backed handoff/forwarding/in-place/dirty/e2e pins plus the new no-list-membership and wrapperless fail-closed pins) joins the existing contributor-lane suite",
|
||||
|
|
@ -224,5 +223,5 @@ BYTE_BASELINE_DEBT = {
|
|||
|
||||
BYTE_DEBT = {
|
||||
"tests/test_devtools_benchmarks.py": 327883,
|
||||
"web/modules/chat.js": 224244,
|
||||
"web/modules/chat.js": 210935,
|
||||
}
|
||||
|
|
|
|||
|
|
@ -137,7 +137,7 @@ class LoadedSkill:
|
|||
from ouroboros.tools.skill_exec import _resolve_runtime_binary, _resolve_script_path
|
||||
|
||||
runtime = (self.manifest.runtime or "").strip().lower()
|
||||
if _resolve_runtime_binary(runtime) is None:
|
||||
if _resolve_runtime_binary(runtime)[0] is None:
|
||||
return False
|
||||
for entry in self.manifest.scripts or []:
|
||||
if not isinstance(entry, dict):
|
||||
|
|
@ -1446,7 +1446,7 @@ def summarize_skills(drive_root: pathlib.Path) -> Dict[str, Any]:
|
|||
available = blocked_by_grants = pending_review = blocker_review = warning_review = broken = 0
|
||||
for s in skills:
|
||||
stale = s.review.is_stale_for(s.content_hash)
|
||||
gate = skill_review_gate(s.review.status, stale=stale)
|
||||
gate = skill_review_gate(s.review.status, stale=stale, findings=s.review.findings)
|
||||
if s.identity_collision:
|
||||
# Readiness probes include lifecycle/dependency state. A collision
|
||||
# has no unique lifecycle identity, so its UI projection must stay
|
||||
|
|
|
|||
|
|
@ -24,6 +24,7 @@ from ouroboros.skill_loader import (
|
|||
SkillReviewState,
|
||||
compute_content_hash,
|
||||
find_skill,
|
||||
load_review_state,
|
||||
save_review_state,
|
||||
skill_state_dir,
|
||||
)
|
||||
|
|
@ -41,10 +42,18 @@ def run_owner_attestation(ctx: Any, drive_root: pathlib.Path, skill: Any, conten
|
|||
PASS finding so the status serializes) and drop the owner-issued marker that
|
||||
load_review_state requires for the verdict to stay valid. Owner-only (the endpoint gates
|
||||
it); the agent can never forge the marker (it is an owner-state file)."""
|
||||
# persist=False: a FAILED attestation attempt must not clobber the skill's existing
|
||||
# review state — the endpoint surfaces the preflight failure (409) and leaves state as-is.
|
||||
# A FAILED attestation preflight persists as a normal review result (so the gate's
|
||||
# fresh ``preflight_failed`` fact and the Repair affordance appear) — but ONLY when
|
||||
# persisting cannot clobber a FRESH valid verdict: review.json absent, or its recorded
|
||||
# content_hash stale for the current payload bytes. A stale CLEAN is knowingly
|
||||
# superseded by a FAIL at the current hash (the fact about the CURRENT bytes wins;
|
||||
# the prior verdict stays in review_history.jsonl). With a fresh verdict on file the
|
||||
# failure still surfaces only through the endpoint (409) and leaves state as-is.
|
||||
persist_failed_preflight = load_review_state(
|
||||
drive_root, skill.name, skill_dir=getattr(skill, "skill_dir", None),
|
||||
).is_stale_for(content_hash)
|
||||
preflight_outcome = _sr._run_deterministic_preflight(
|
||||
ctx, drive_root, skill, content_hash, persist=False,
|
||||
ctx, drive_root, skill, content_hash, persist=persist_failed_preflight,
|
||||
binding=getattr(ctx, "_skill_review_resolved_binding", None),
|
||||
)
|
||||
if preflight_outcome is not None:
|
||||
|
|
|
|||
|
|
@ -847,7 +847,7 @@ def _outcome_payload(
|
|||
job_data: Dict[str, Any] | None = None,
|
||||
) -> Dict[str, Any]:
|
||||
status = normalize_skill_review_status(outcome.status)
|
||||
gate = skill_review_gate(status)
|
||||
gate = skill_review_gate(status, findings=outcome.findings)
|
||||
payload: Dict[str, Any] = {
|
||||
"skill": outcome.skill_name,
|
||||
"status": status,
|
||||
|
|
|
|||
|
|
@ -88,12 +88,8 @@ def aggregate_skill_review_status(
|
|||
# as STATUS_PENDING (non-executable under every enforcement mode) and MUST
|
||||
# reload that way — never as advisory-overridable BLOCKERS or WARNINGS —
|
||||
# so honor them before the severity-driven aggregation below.
|
||||
for finding in findings:
|
||||
if finding.get("verdict") == "FAIL" and (
|
||||
finding.get("item") in ("skill_preflight", "plugin_api_admission")
|
||||
or str(finding.get("model") or "") in ("deterministic_preflight", "plugin_api_admission")
|
||||
):
|
||||
return STATUS_PENDING
|
||||
if preflight_failed(findings):
|
||||
return STATUS_PENDING
|
||||
is_official_hub = review_profile == "official_hub"
|
||||
has_critical_fail = False
|
||||
has_warning_fail = False
|
||||
|
|
@ -177,13 +173,48 @@ def count_trailing_warnings_rounds(
|
|||
return count
|
||||
|
||||
|
||||
def skill_review_gate(status: str, *, stale: bool = False, enforcement: Optional[str] = None) -> Dict[str, Any]:
|
||||
def preflight_failed(findings: Any) -> bool:
|
||||
"""True when the persisted findings carry a deterministic preflight FAIL.
|
||||
|
||||
The SSOT for the shape is the finding `_run_deterministic_preflight`
|
||||
persists (item=skill_preflight / model=deterministic_preflight) — plus the
|
||||
ABI-1 PluginAPI admission refusal (item/model=plugin_api_admission), the
|
||||
same structural-gate class; the same condition drives the pending
|
||||
aggregation above. UI cards use this fact to offer Repair instead of a
|
||||
Re-review that would deterministically fail again (#335).
|
||||
"""
|
||||
for finding in findings or []:
|
||||
if not isinstance(finding, dict):
|
||||
continue
|
||||
if finding.get("verdict") == "FAIL" and (
|
||||
finding.get("item") in ("skill_preflight", "plugin_api_admission")
|
||||
or str(finding.get("model") or "") in ("deterministic_preflight", "plugin_api_admission")
|
||||
):
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def skill_review_gate(
|
||||
status: str, *, stale: bool = False, enforcement: Optional[str] = None,
|
||||
findings: Any = None,
|
||||
) -> Dict[str, Any]:
|
||||
"""Structured, agent-facing explanation of whether a review is executable.
|
||||
|
||||
Deterministic hard-gate failures (e.g. skill_preflight) are persisted as
|
||||
STATUS_PENDING by `_run_deterministic_preflight`, so they are non-executable
|
||||
here under every enforcement mode without needing per-caller findings — only
|
||||
LLM blocker verdicts are overridable by advisory enforcement.
|
||||
|
||||
``findings`` is optional: a caller that has the persisted findings gets a
|
||||
``preflight_failed`` key in the gate (the typed fact behind the Repair
|
||||
affordance, #335) plus its companion ``preflight_failed_stale``. Absence
|
||||
of the keys means the caller could not know — they are never fabricated.
|
||||
A STALE review's persisted failure no longer describes the current payload
|
||||
bytes (the owner may have fixed it by hand), so ``preflight_failed`` is
|
||||
True only while the findings are fresh; the recorded-but-stale failure
|
||||
surfaces as ``preflight_failed_stale`` instead, and the card offers BOTH
|
||||
actions: the cheap Re-review (which reruns the preflight) stays primary,
|
||||
with Repair offered based on the last recorded preflight.
|
||||
"""
|
||||
raw_status = normalize_skill_review_status(status)
|
||||
if enforcement is None:
|
||||
|
|
@ -229,4 +260,8 @@ def skill_review_gate(status: str, *, stale: bool = False, enforcement: Optional
|
|||
"blocking_reason": reason,
|
||||
"review_enforcement": enforcement,
|
||||
"summary": summary,
|
||||
**({
|
||||
"preflight_failed": (not stale) and preflight_failed(findings),
|
||||
"preflight_failed_stale": bool(stale) and preflight_failed(findings),
|
||||
} if findings is not None else {}),
|
||||
}
|
||||
|
|
|
|||
|
|
@ -149,6 +149,45 @@ def apply_task_start_settings() -> None:
|
|||
apply_settings_to_env(effective)
|
||||
|
||||
|
||||
def apply_task_start_settings_or_disclose(task_id: str, emit_live_log: Any) -> None:
|
||||
"""Task-start settings reload with a LOUD failure path (#285).
|
||||
|
||||
A silent failure breaks the save-time promise "the saved changes apply
|
||||
from the next task": the task would run on the previously applied
|
||||
configuration with nobody told. The task itself stays runnable
|
||||
(fail-open), but the breakage becomes a visible live-log fact.
|
||||
|
||||
The common corruption case is probed explicitly: ``load_settings`` falls
|
||||
back to defaults+env on an unreadable or malformed settings.json instead
|
||||
of raising, which would keep exactly the silence this wrapper exists to
|
||||
break. A MISSING file is legitimate (defaults-only install), not a fault.
|
||||
"""
|
||||
try:
|
||||
from ouroboros import config as _config
|
||||
|
||||
try:
|
||||
raw_settings_text = _config.SETTINGS_PATH.read_text(encoding="utf-8")
|
||||
except FileNotFoundError:
|
||||
raw_settings_text = None
|
||||
if raw_settings_text is not None:
|
||||
json.loads(raw_settings_text)
|
||||
apply_task_start_settings()
|
||||
except Exception as exc:
|
||||
import logging
|
||||
|
||||
logging.getLogger(__name__).error(
|
||||
"Task-start settings reload failed; this task runs on the previously applied configuration",
|
||||
exc_info=True,
|
||||
)
|
||||
emit_live_log(
|
||||
"task_start_settings_reload_failed",
|
||||
task_id=task_id,
|
||||
error=f"{type(exc).__name__}: {exc}",
|
||||
message=("Settings reload failed at task start: this task runs "
|
||||
"on the previously applied configuration."),
|
||||
)
|
||||
|
||||
|
||||
def _resolution(
|
||||
settings: Mapping[str, Any], *, allow_undecided_legacy: bool,
|
||||
) -> ConfiguredSubagentsResolution:
|
||||
|
|
@ -558,11 +597,17 @@ def prepare_delegate_start_actor(
|
|||
drive_root, str(getattr(ctx, "task_id", "") or "")
|
||||
)
|
||||
if any(blockers.values()):
|
||||
live_ids = [str(v) for key in ("open_run_ids", "pending_invocation_ids",
|
||||
"undisposed_patch_run_ids")
|
||||
for v in (blockers.get(key) or [])]
|
||||
shown = ", ".join(live_ids[:4]) + (
|
||||
f" (+{len(live_ids) - 4} more)" if len(live_ids) > 4 else "")
|
||||
return {}, _fail(
|
||||
"delegate_start", "replacement_requires_settlement",
|
||||
"This task still owns an unsettled run/start invocation or an undisposed "
|
||||
"captured patch. Prove the predecessor absent or terminal and explicitly "
|
||||
"dispose any captured patch before starting a replacement.",
|
||||
f"captured patch ({shown}). Wait for or cancel an open run; replay a "
|
||||
"pending invocation with retry_of=<invocation id>; dispose a captured "
|
||||
"patch explicitly (#364).",
|
||||
**blockers,
|
||||
)
|
||||
return {
|
||||
|
|
|
|||
|
|
@ -150,10 +150,11 @@ def delegated_execution_workspace_root(
|
|||
gateway: Any, shape: DelegatedRunShape, root: str,
|
||||
) -> str:
|
||||
"""Return the live-workspace field only when the engine's strict schema accepts it."""
|
||||
from ouroboros.config import CLAUDEXOR_DELEGATED_WORKSPACE_ROOT_MIN_VERSION
|
||||
from ouroboros.gateways.claudexor import engine_at_least
|
||||
|
||||
version = str(getattr(gateway, "engine_version", "") or "")
|
||||
supported = engine_at_least(version, "3.8.1")
|
||||
supported = engine_at_least(version, CLAUDEXOR_DELEGATED_WORKSPACE_ROOT_MIN_VERSION)
|
||||
return str(root) if supported and shape.delegated and shape.isolation == "live" else ""
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -68,6 +68,22 @@ def send_provider_death_notice(
|
|||
return True
|
||||
|
||||
|
||||
def stamp_root_final_phase(send_event: Dict[str, Any], task: Dict[str, Any], *, post_task_open: bool) -> None:
|
||||
"""Type a root's final frame for the client's live conclusion gate.
|
||||
|
||||
With post-task synthesis still OPEN the owner's answer leaves early: the
|
||||
typed phase marker (progress_meta merges into the WS chat payload) holds
|
||||
the card on "Finalizing…" until the settled task_done, instead of the
|
||||
early final reading as the task's terminal conclusion. With post-task
|
||||
already settled a DIRECT turn's bare final IS the turn's terminal word
|
||||
(#369) — managed roots keep their task_done conclusion untouched.
|
||||
"""
|
||||
if post_task_open:
|
||||
send_event.setdefault("progress_meta", {})["task_phase"] = "finalizing"
|
||||
elif task.get("_is_direct_chat"):
|
||||
send_event.setdefault("progress_meta", {})["task_terminal_status"] = "completed"
|
||||
|
||||
|
||||
def prepare_terminal_send_event(
|
||||
env_drive_root: Any, task: Dict[str, Any], text: str,
|
||||
usage: Dict[str, Any], send_event: Dict[str, Any],
|
||||
|
|
@ -75,6 +91,13 @@ def prepare_terminal_send_event(
|
|||
) -> Dict[str, Any]:
|
||||
"""Preserve raw host salvage, then build the one live/replay projection."""
|
||||
origin = str(usage.get("terminal_origin") or "")
|
||||
if ephemeral and not presence:
|
||||
# #369: an ephemeral decision's task_done frame is dropped at the
|
||||
# client's log-event entry by design, so this final is the turn's
|
||||
# ONLY conclusion vehicle. The typed fact mirrors the direct-error
|
||||
# branch (supervisor/workers.py stamps task_terminal_status="failed")
|
||||
# and lets the live concludesTurn gate settle the activity.
|
||||
send_event.setdefault("progress_meta", {})["task_terminal_status"] = "completed"
|
||||
if ephemeral or presence or origin not in {
|
||||
TERMINAL_ORIGIN_MODEL_FINAL, TERMINAL_ORIGIN_HOST_SALVAGE,
|
||||
}:
|
||||
|
|
@ -119,11 +142,13 @@ def deliver_final_message_live(
|
|||
) -> bool:
|
||||
"""Send the buffered FINAL ``send_message`` through the live worker queue.
|
||||
|
||||
The buffer can also hold proactive ``send_user_message`` events queued
|
||||
mid-task (they carry no ``task_id``), so the final answer is selected by
|
||||
the finalizing task's id — falling back to the LAST send_message — never
|
||||
the first match, which would ship a proactive text early while the answer
|
||||
stayed hostage to blocking post-task.
|
||||
The buffer can also hold proactive ``send_user_message`` events that fell
|
||||
back to deferred delivery mid-task (live-first frames stamp ``task_id``
|
||||
too), so the final answer is selected as the LAST send_message matching
|
||||
the finalizing task's id — the host appends the terminal frame after all
|
||||
tool-time frames, so it wins the last-match scan — never the first match,
|
||||
which would ship a proactive text early while the answer stayed hostage
|
||||
to blocking post-task.
|
||||
|
||||
Never lost, never doubled — without treating ``queue.put()`` as a delivery
|
||||
receipt: the buffered copy is KEPT (an event can still die between put and
|
||||
|
|
|
|||
|
|
@ -126,8 +126,27 @@ def _root_task_acceptance_review_cap(
|
|||
)
|
||||
|
||||
|
||||
def _claim_for_paid_identity(claims: Any, paid_identity: str) -> Optional[Dict[str, Any]]:
|
||||
"""The claim row this tree already bought for one A-material paid identity.
|
||||
|
||||
ONE answer shared by the free-refusal projection and the atomic claim, so the
|
||||
dispatch seam can never refuse what the wallet would have allowed (or vice
|
||||
versa). An empty identity matches nothing: a pre-A-material row keys on its
|
||||
binding hash alone."""
|
||||
identity = str(paid_identity or "").strip().lower()
|
||||
if not identity:
|
||||
return None
|
||||
return next(
|
||||
(
|
||||
row for row in (claims or {}).values()
|
||||
if isinstance(row, dict) and str(row.get("paid_identity") or "") == identity
|
||||
),
|
||||
None,
|
||||
)
|
||||
|
||||
|
||||
def project_task_acceptance_review_capacity(
|
||||
ctx: Any, *, binding_hash: str = "", task_id: str = "",
|
||||
ctx: Any, *, binding_hash: str = "", task_id: str = "", paid_identity: str = "",
|
||||
) -> Dict[str, Any]:
|
||||
"""Read the canonical root's paid acceptance-wallet projection.
|
||||
|
||||
|
|
@ -186,6 +205,12 @@ def project_task_acceptance_review_capacity(
|
|||
claimed = len(claims)
|
||||
remaining = None if cap is None else max(0, cap - claimed)
|
||||
requested_binding = str(binding_hash or "").strip().lower()
|
||||
# Seen = this dispatch was already PAID for, under either identity: the
|
||||
# exact binding (as before) or the A-material paid identity, so a resubmit
|
||||
# that only moved the evidence revision cannot buy a second panel.
|
||||
binding_seen = bool(requested_binding and requested_binding in claims) or (
|
||||
_claim_for_paid_identity(claims, paid_identity) is not None
|
||||
)
|
||||
projection = {
|
||||
**base,
|
||||
"state": "available",
|
||||
|
|
@ -193,7 +218,7 @@ def project_task_acceptance_review_capacity(
|
|||
"cap_cycles": cap,
|
||||
"claimed_cycles": claimed,
|
||||
"remaining_cycles": remaining,
|
||||
"binding_seen": bool(requested_binding and requested_binding in claims),
|
||||
"binding_seen": binding_seen,
|
||||
}
|
||||
try:
|
||||
from ouroboros.cancel_intents import cancel_pending
|
||||
|
|
@ -299,6 +324,11 @@ _TASK_ACCEPTANCE_REVIEW_CLAIM_FIELDS = frozenset({
|
|||
"binding_hash", "candidate_hash", "evidence_revision", "fence_hash",
|
||||
"claimed_at", "claimed_by_task_id",
|
||||
})
|
||||
# A-material (2026-08-30), additive: the identity the panel was actually PAID
|
||||
# for — candidate answer plus new obligation dispositions, evidence revision
|
||||
# deliberately excluded. Rows written before it carry no such key and stay valid;
|
||||
# new rows carry both, so the binding-keyed reads above never change meaning.
|
||||
_TASK_ACCEPTANCE_REVIEW_CLAIM_OPTIONAL_FIELDS = frozenset({"paid_identity"})
|
||||
|
||||
|
||||
def _empty_task_acceptance_review_state(root_task_id: str) -> Dict[str, Any]:
|
||||
|
|
@ -326,7 +356,11 @@ def _validated_task_acceptance_review_state(
|
|||
for binding_hash, claim in claims.items():
|
||||
if not _PLAN_REVIEW_HASH_RE.fullmatch(str(binding_hash or "")):
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: claim key is invalid")
|
||||
if not isinstance(claim, dict) or set(claim) != _TASK_ACCEPTANCE_REVIEW_CLAIM_FIELDS:
|
||||
if not isinstance(claim, dict) or not (
|
||||
_TASK_ACCEPTANCE_REVIEW_CLAIM_FIELDS
|
||||
<= set(claim)
|
||||
<= (_TASK_ACCEPTANCE_REVIEW_CLAIM_FIELDS | _TASK_ACCEPTANCE_REVIEW_CLAIM_OPTIONAL_FIELDS)
|
||||
):
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: claim shape is invalid")
|
||||
if str(claim.get("binding_hash") or "") != str(binding_hash):
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: claim identity mismatch")
|
||||
|
|
@ -335,6 +369,12 @@ def _validated_task_acceptance_review_state(
|
|||
raise ValueError(
|
||||
f"TASK_ACCEPTANCE_REVIEW_STATE_INVALID: {key} is invalid"
|
||||
)
|
||||
if "paid_identity" in claim and not _PLAN_REVIEW_HASH_RE.fullmatch(
|
||||
str(claim.get("paid_identity") or "")
|
||||
):
|
||||
raise ValueError(
|
||||
"TASK_ACCEPTANCE_REVIEW_STATE_INVALID: paid_identity is invalid"
|
||||
)
|
||||
expected_binding = review_binding_hash(
|
||||
candidate_hash=str(claim["candidate_hash"]),
|
||||
evidence_revision=str(claim["evidence_revision"]),
|
||||
|
|
@ -479,6 +519,13 @@ def claim_task_acceptance_review_cycle(
|
|||
if binding_fields["binding_hash"] != expected_binding:
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: binding digest mismatch")
|
||||
binding = binding_fields["binding_hash"]
|
||||
# A-material: the identity the panel is PAID for. Absent (a pre-A-material
|
||||
# caller) => the binding hash keeps being the paid identity, i.e. exactly the
|
||||
# old behaviour; present => it also refuses a second claim for the same
|
||||
# material under a different binding.
|
||||
paid_identity = str((review_binding or {}).get("paid_identity") or "").strip().lower()
|
||||
if paid_identity and not _PLAN_REVIEW_HASH_RE.fullmatch(paid_identity):
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: paid identity is invalid")
|
||||
claimant = str(claimed_by_task_id or "").strip()
|
||||
if not claimant:
|
||||
raise ValueError("TASK_ACCEPTANCE_REVIEW_STATE_INVALID: claimant is absent")
|
||||
|
|
@ -499,7 +546,7 @@ def claim_task_acceptance_review_cycle(
|
|||
cap = _root_task_acceptance_review_cap(root_result)
|
||||
resolved["cap"] = cap
|
||||
claims = dict(state.get("claims_by_binding") or {})
|
||||
prior = claims.get(binding)
|
||||
prior = claims.get(binding) or _claim_for_paid_identity(claims, paid_identity)
|
||||
if prior is not None:
|
||||
decision.update({
|
||||
"status": "unknown",
|
||||
|
|
@ -511,6 +558,7 @@ def claim_task_acceptance_review_cycle(
|
|||
return None
|
||||
claims[binding] = {
|
||||
**binding_fields,
|
||||
**({"paid_identity": paid_identity} if paid_identity else {}),
|
||||
"claimed_at": utc_now_iso(),
|
||||
"claimed_by_task_id": claimant,
|
||||
}
|
||||
|
|
@ -532,6 +580,7 @@ def claim_task_acceptance_review_cycle(
|
|||
return {
|
||||
**decision,
|
||||
"binding_hash": binding,
|
||||
"paid_identity": paid_identity,
|
||||
"cycles_paid": paid,
|
||||
"max_cycles": cap,
|
||||
"remaining_cycles": None if cap is None else max(0, cap - paid),
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Add a link
Reference in a new issue