Commit graph

2469 commits

Author SHA1 Message Date
Ouroboros
6a2ee44ac2 docs/tests: narrow the transport-death repeat wording to Host-executed effects; state the 409 condition exactly
Reviewer P2: a lost response proves no Host-executed tool or correspondent action from the failed attempt, not that nothing ran anywhere — priced read-only provider-owned retrieval (server-side web search) may rerun on the repeat. The 409 and owner notice apply when the round ends unresolved without a further permitted repeat (deadline or a finalize control can refuse the first repeat). Test comment on the owner-notice writer corrected; chapter budgets untouched.
2026-09-26 02:00:57 +03:00
Ouroboros
eb4e42b972 test/docs: pin the same-round transport-death repeat as a no-effect attempt for Presence; unknown-outcome notice test on the 409 contract
Regression drives the real round dispatcher with a ledger-backed fake provider: one dispatched ReadError, one repeat, two attempts, no tool run or send between them, no forced final, no presence_unknown_outcome. test_provider_terminal_notice converted from the pre-TZ3 200-deferred expectation to the 409 presence_attempt_outcome_unknown contract with admitted child custody preserved. Ch.06/ch.12 state the contract.
2026-09-26 01:26:42 +03:00
Ouroboros
013289a4ca fix: Presence guard refuses on provenance, not the infra_failed axis (Astra/Fable delta findings)
presence_runner._terminal_refusal keys on a missing canonical terminal or the new durable presence_unknown_outcome marker (stamped by agent_task_pipeline from the loop no-resend predicate); confirmed outages and overflows keep the TZ2 deferred projection with admitted child custody. Regressions drive the real emit_task_results pipeline over host-fallback, salvaged-draft and death-record shapes.
2026-09-26 00:38:18 +03:00
Ouroboros
54c2be5517 fix: refuse Presence acknowledgement on infrastructure terminals and lost terminal writes
Independent exact-range review (Astra, 55ad78213..15dbc7533) found two P1 paths
where the Host acknowledged a Presence event the model never answered:

- a quota-refused primary whose fallback died with an unknown outcome is worded
  provider_unavailable by the forced rail (the unknown fence outranks the
  refusal source), so a guard keyed on resource_refusal_no_resend alone let the
  host-authored terminal project to completed/silent (HTTP 200) live and on
  replay, and the adapter could drop the event;
- a terminal write that fails after the durable start barrier is swallowed by
  the pipeline, and the in-memory presence_result envelope was returned as the
  answer while the durable row still said RUNNING.

One helper, _terminal_refusal, now owns the verdict the durable row demands
before acknowledgement: a confirmed resource refusal keeps the event
(presence_resources_unavailable); no canonical terminal (RUNNING/INTERRUPTED,
reconciled placeholder, lost terminal write) or any infra_failed execution
axis is presence_attempt_outcome_unknown — 409, empty external speech, the
admitted work_ref preserved (durable metadata, else the execution's own
handoff fact), one serialized owner-only notice, never a retry certificate.
Both guards (live and cached replay) consume it. Two consumer regressions
through the real Host app cover both paths; test fakes that returned an
envelope without writing a terminal now write the durable row the real
pipeline writes. Chapter 12 records the rule.
2026-09-25 23:51:52 +03:00
Ouroboros
5507134f71 Merge remote-tracking branch 'origin/ouroboros' (55ad78213) into ouroboros-agent/presence-resilience
Semantic integration of #1299 (memory history visibility) and #1207 (reasoning-effort descriptor + Z.ai provider). One textual conflict: tests/test_persistence_inventory.py EXPECTED_SCAN_PATHS resolved 295 (+1 Presence quarantine members, -1 removed memory journal rewrite path); verified by test_persistence_inventory + reference-book/inventory/domain checks (48 passed).
2026-09-25 23:16:14 +03:00
Ouroboros
49e509b866 fix: rebind primary quota waits and serialize Presence recovery notices 2026-09-25 20:43:28 +03:00
Ouroboros
47c96ed29e fix: bound Presence preparation and retain fallback quota waits 2026-09-25 19:38:26 +03:00
Ouroboros
6830b44871 docs: describe Presence retry without retrospective residue 2026-09-25 18:37:10 +03:00
Ouroboros
80c278cb53 Merge current ouroboros target (6fe01dadf) into pr/reasoning-effort
Regenerated docs/inventories/DATA_LAYOUT_INVENTORY.md on the combined tree (both
sides had regenerated it).
2026-09-25 18:21:15 +03:00
Ouroboros
d3d7592849 Close the round-2 review findings on the Z.ai provider
Four independent reviewers (Fable 5.1, Opus 5.5, Codex gpt-6-sol, Grok 4.7)
ran the Ouroboros scope-review brief on the reworked candidate; all four
returned "merge after small fixes" and no round-1 blocker survived. The
accepted findings, each mirroring the MiniMax region pattern:

- The one-click Colab collector never asked for ZAI_API_KEY, so a Z.ai-only
  notebook was prompted for OpenRouter (colab_bootstrap.provider_keys).
- ZAI_PLAN is a transport choice, not a credential: it no longer satisfies the
  onboarding "has a provider" checks (settings_setup_contract, the wizard's
  two lists), and an unknown plan value is refused at settings save and by the
  provider Test instead of silently selecting pay-as-you-go.
- The Capability Evidence route readers (gateway/settings._active_main_route,
  reviewer_window.reviewer_route) resolve the plan's endpoint, so the two
  Z.ai plans no longer share one route fingerprint.
- The task loop classifies Z.ai's HTTP 429 code 1113 as quota_exhausted
  (typed code, exact match) instead of retrying it as a transient rate limit.
- OpenRouter's GLM namespace is z-ai/, so the catalog label key follows it.
- Docs: Z.ai joins the exclusive-direct-provider list in 02-startup-onboarding
  and the forbidden skill settings list in CREATING_SKILLS; the 02-naming
  sentence names the two provider-specific projections instead of implying
  direct OpenAI carries no effort; glm ids join the onboarding suggestions.

Declined as disproportionate or out of scope, with the reason recorded in
the review ledger: the CHECKLISTS.md prose parenthetical (protected file;
the runtime deny list is authoritative), the GAIA/Terminal-Bench launcher
key lists (benchmark-only), a recorded live wire fixture (no key), and
GLM-5.2's skip-thinking on none/minimal (owner-accepted, disclosed in the
external-fact inventory).
2026-09-25 18:19:15 +03:00
Ouroboros
bcd9455e1d fix: retain Presence resource refusals and admitted work across failures 2026-09-25 18:10:16 +03:00
Ouroboros
b463bb3d93 Add the Z.ai (GLM) direct provider with effort projection at the send boundary
zai:: joins the direct providers exactly the way deepseek:: did: prefix and
credential registry, ZAI_API_KEY plus a ZAI_PLAN endpoint selector
(provider_models.resolve_zai_base_url: empty/payg = api.z.ai/api/paas/v4,
coding = the Coding Plan endpoint), the routing target, live catalog fetch,
provider Test, settings card, onboarding contract, review-fallback roles,
single-provider startup and review detection, secret masking, benchmark
env hygiene, and docs.

Reasoning effort now reaches Z.ai. The provider serves an ABSENT
reasoning_effort at its maximum tier, so every call on the old generic
compatible route was billed at max regardless of the configured effort.
The canonical scale is projected onto Z.ai's own low/high/max enum
(ZAI_REASONING_EFFORT_ALIASES: none/minimal -> low, medium -> high,
xhigh/ultra -> max), disclosed as reasoning_effort_clamped when the tier
changes; GLM-5.3 rejects every other value and cannot disable thinking
(HTTP 400 code 1210), and forced tool_choice works with thinking on, so
there is no DeepSeek-style suppression arm. The projection is keyed on the
provider id the owner configured, never on a model name: a GLM served
from an owner's own OpenAI-compatible endpoint keeps today's behavior.

The provider port is the contributor's own work from the closed PR #1194,
narrowed to Z.ai (the DashScope and Moonshot lanes were not measured and
stay out). 07-configuration gains two settings rows and one route
paragraph (budget 37300 -> 38400), 02-naming records the dated Z.ai
probe in the external-fact inventory, and the onboarding bootstrap
fixture and data-layout inventory are regenerated.

Co-authored-by: josephsteuerjr <josephsteuerjr@gmail.com>
2026-09-25 17:41:54 +03:00
Ouroboros
407868969a Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/presence-resilience
# Conflicts:
#	docs/architecture/12-host-service-companions-and-chat-ids.md
#	tests/test_reference_book_budgets.py
2026-09-25 17:39:07 +03:00
Ouroboros
fa741a7fe9 Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/tz3-history-visibility 2026-09-25 17:24:18 +03:00
Ouroboros
b60b8fe98a Narrow the reasoning-effort change to its landed core
Revert the route-descriptor registry, the descriptor-driven effort pickers,
the Anthropic builder indirection and the effort_not_carried plumbing back
to the target's state, and restore DeepSeek's documented projection
(xhigh -> high per api-docs.deepseek.com Thinking Mode).

What the maintainers found on the combined tree: nine target-green tests
went red (five llm_golden compatible/cloudru/minimax cases, three
host-message-metadata cases, the provider_models/config leaf guard), the
07-configuration chapter left its byte budget, the picker hid saved and
working tiers (xhigh/max/none) and matched the catalog by bare model id,
the reviewer-slot wiring passed a fourth argument that the local wrapper
dropped, and the new usage fact never reached events.jsonl. The z.ai GLM
effort carriage itself is kept and lands as a first-class provider in the
next commit.

Kept from this change: Z.ai's HTTP 429 code 1113 ("Insufficient balance")
is a billing fact, so the provider Test maps it to "No credits" instead of
"Rate limited" (typed code only; the status copy stays provider-neutral).
2026-09-25 17:18:15 +03:00
Ouroboros
5f1b848570 fix: preserve quota refusal and retry successor provenance 2026-09-25 17:03:45 +03:00
Ouroboros
f21425cd8b Merge current ouroboros target (86d4e6298) into pr/reasoning-effort 2026-09-25 16:48:45 +03:00
Ouroboros
657005f68e fix: retain Main context on unreadable dialogue metadata 2026-09-25 16:39:09 +03:00
Ouroboros
2933fa93db Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/presence-resilience 2026-09-25 16:29:09 +03:00
Ouroboros
925705bb09 fix: retry Presence events only after proven no-effect quota refusal 2026-09-25 16:28:49 +03:00
Ouroboros
9868f8890f docs: distinguish default Presence binding scope from selected global readers 2026-09-25 15:28:09 +03:00
Ouroboros
6675073bee Merge remote-tracking branch 'origin/ouroboros' into ouroboros/tz2-coordination 2026-09-25 14:40:24 +03:00
Ouroboros
b2b5b4f5a3 feat: preserve memory history and expose pending nominations 2026-09-25 14:30:46 +03:00
Ouroboros
7d8e742d12 wip: preserve Presence TZ2 corrections pending Q4 and full review 2026-09-25 09:37:03 +03:00
Ouroboros
ec1dfa0685 Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/tz3-memory 2026-09-25 09:25:09 +03:00
Ouroboros
8fc087c275 fix: bound knowledge edit receipts and preserve cyclic YAML 2026-09-25 09:06:15 +03:00
Ouroboros
dcdd9984fb fix: retain malformed-note append under truthful heading delta 2026-09-25 08:43:11 +03:00
Ouroboros
7ed80c8a79 Merge remote-tracking branch 'origin/ouroboros' into ouroboros/tz2-coordination
# Conflicts:
#	tests/test_reference_book_budgets.py
2026-09-25 07:38:33 +03:00
Ouroboros
3d941e7feb Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/tz3-memory 2026-09-25 07:35:48 +03:00
Ouroboros
26aef771b4 fix: bind Cowork evaluator score to run protocol provenance 2026-09-25 07:09:54 +03:00
Ouroboros
4d16911126 feat: add exact source-bound knowledge edits
Expose a body-only edit mode with one exact anchor and revision CAS. Preserve frontmatter, capture the complete old/new source, and report an honest host-derived delta without selecting the semantic writer. Add consumer and refusal regressions; keep release carriers unchanged for contributor integration.
2026-09-25 06:45:34 +03:00
Ouroboros
09b4682521 fix: preserve Presence binding authority in delegated children
Keep the host-derived own-binding fence across scheduling and supervisor steering without turning children into Presence speakers. Add real scheduling and fake-model consumer regressions.
2026-09-25 06:25:08 +03:00
Ouroboros
f7890e7659 fix: bind Presence retries to durable starts and source events
Keep Host auth work from saturating control threads, refuse uncertain JSONL boundaries, and retain quota-exhausted events without speaking draft replies. This is an intermediate contributor commit pending the remaining no-effect retry decision and exact-diff review.
2026-09-25 06:12:00 +03:00
Ouroboros
720fd98152 feat: add claimed Cowork evaluation and opt-in residual audit 2026-09-25 06:00:18 +03:00
Ouroboros
0866ee9ba5 docs: keep Presence lifecycle explanation inside chapter budget 2026-09-25 05:46:15 +03:00
Ouroboros
6bc2013bdb Merge remote-tracking branch 'origin/ouroboros' into ouroboros/tz2-coordination
# Conflicts:
#	tests/test_reference_book_budgets.py
2026-09-25 05:31:49 +03:00
Ouroboros
8836d20fa5 docs: fit tool-discovery map after upstream integration 2026-09-25 05:24:48 +03:00
Ouroboros
a4ca3c4e56 Merge remote-tracking branch 'origin/ouroboros' into fix/1262-tool-name-recovery
# Conflicts:
#	docs/inventories/DATA_LAYOUT_INVENTORY.md
#	tests/test_reference_book_budgets.py
2026-09-25 05:18:28 +03:00
Ouroboros
6fc0afc268 fix: keep MCP revocation effective across cached catalog hits
Recheck saved authority before Safety only for exact hits, keep misses pure, and confine catalog omissions to callable names. Bound ambiguous identity guidance and cover resource-blocked builtin discovery.
2026-09-25 04:35:34 +03:00
Ouroboros
4d149ff677 ouroboros: checkpoint after task 4b8cb25784ee4f17 — ТЗ2 — завершить самокоординацию Presence 2026-09-25 04:24:24 +03:00
Ouroboros
b299367135 ouroboros: checkpoint after task 410dfca20a30497b — Presence ТЗ3 — завершить сохранённый кандидат 2026-09-25 04:08:04 +03:00
Ouroboros
3558e14977 Regenerate the data-layout inventory for the target's chapter 01 edit
The moved target (829164b90, via the headless lineage-read PR) edited the
data-layout tree in docs/architecture/01-high-level-architecture.md without
regenerating docs/inventories/DATA_LAYOUT_INVENTORY.md, so
tests/test_generated_inventories.py::test_data_layout_inventory_is_byte_identical
is red on the target itself. Only the source SHA line changes; the entries are
identical. Regenerated with scripts/regenerate_inventories.py.
2026-09-25 04:07:39 +03:00
Ouroboros
fa0bbad666 Merge managed/ouroboros 829164b90 into the #1002/#1230 branch 2026-09-25 03:57:10 +03:00
Ouroboros
5fb287a63d Apply the review wave: slim budget line everywhere, flush before restart, honest batch time budget
Findings from the simulated triad (gpt-6-astra, grok-4.7, fable-5.1), one batch:

- The budget line's unbounded-budget branch still rendered the full
  usage_breakdown; it now reads usage_writer_snapshot like /api/state, so the
  chapter 05 sentence that both read the slim shapes is true on every install.
- A restart that lands right after a drain now flushes that turn's dirty budget
  projection before leaving the loop, so the last drained llm_usage events
  still reach state.json.
- The events batch time budget is checked between handlers; the docstring,
  the constant comment and chapter 05 now say a running handler can overrun it
  instead of promising a hard elapsed-time bound.
- The socket-open project refresh forces its read: the in-flight read predates
  the socket and was just invalidated, so joining it would bring nothing fresh.
- The update_budget_from_usage docstring distinguishes the loop's once-per-turn
  llm_usage path from direct callers; plan labels removed from test docstrings.
2026-09-25 03:56:23 +03:00
Ouroboros
829164b908
Merge PR #1283: retain Cowork evaluator evidence (#1261 A)
fix: retain Cowork evaluator evidence independently of execution (#1261 A)
2026-09-25 03:51:28 +03:00
Ouroboros
fdebeabf94 Merge remote-tracking branch 'origin/ouroboros' into fix/1262-tool-name-recovery 2026-09-25 03:49:46 +03:00
Ouroboros
95078b23bc fix: recover callable tool names on exact catalog misses
Provide scoped MCP raw-to-wire discovery, pure pre-safety name resolution, and a shared policy-filtered refusal path for tool namespaces. Preserve exact dispatch and extension adoption; cover real registry consumers and classification.
2026-09-25 03:46:59 +03:00
Ouroboros
1e8fe5e8c0 Gate the page-wide /api/state reads to one in flight
The chat header's 3 s refresh and the projects-nav refresh both read
/api/state through the shared snapshot sequencer, which ordered responses
but never limited how many reads were in flight. With a slow server answer
the polls piled up (5 concurrent from the loop, 6 with the socket's own SHA
read) and the project panel's history read queued behind them for minutes
(issue #1102 / #1195 class: header polls stacked behind slow /api/state).

createStateSnapshotSequencer gains gate(force): one read in flight page-wide;
a periodic tick during a read starts nothing and resolves once that read
settles; forced refreshes that land mid-flight coalesce into exactly one
follow-up read that starts after the in-flight one settles (success, HTTP
error, thrown fetch), and every forced caller resolves after it. begin()
and the generation ordering are unchanged; a gated request settles through
apply/fail of the same request object. chat.js and app.js read through the
gate; the boot prefetch, socket-open nav refresh and 20 s nav poll join an
in-flight read instead of forcing a second one. Test doubles learn gate().
2026-09-25 03:26:50 +03:00
Ouroboros
6ce5df9542 Bound the supervisor events pass and write the budget projection once per turn
Every llm_usage worker event made the supervisor loop render the FULL usage
breakdown (five grouped axes plus the per-root projection over the whole
ledger) to persist a handful of totals, take STATE_LOCK and rewrite
state.json, and the loop drained the event queue to empty before it read the
owner's bridge; at one settled attempt per second the events phase never
returned to intake and the chat went silent while the header said Online.

The compatibility writer now reads a slim render-cached snapshot
(usage_writer_snapshot: totals, the ordering marker, the OpenRouter bucket its
drift check compares, the totals-only projection) computed from the same
validated rows as usage_breakdown, so the marker and the money still come from
one read; /api/state and the budget line read the same slim shapes.
_projection_from_final moves verbatim into the _usage_rows leaf and is
re-exported. An llm_usage event only marks the loop context dirty (its row and
live frame say "deferred"); the loop flushes one write per turn AFTER bridge
intake, keeps the flag on a refused or failed write and retries no more often
than BUDGET_PROJECTION_RETRY_SEC, and the OpenRouter ground-truth check fires on
crossing each multiple of 50 physical calls (a coalesced write may jump 49 to
51). One events pass drains at most SUPERVISOR_EVENT_BATCH_MAX_EVENTS or
SUPERVISOR_EVENT_BATCH_MAX_SEC, FIFO, restart requests and event-lag
observation preserved, the remainder next turn; a turn that hit its bound skips
the idle sleep. The drain and the flush live in server_liveness beside the
other loop-phase helpers, so server.py shrinks.

Tests pin the snapshot against the full breakdown on every key the writer reads
with the full render spied out, byte-identical state.json from either render,
the 49 to 51 crossing, N events to one write after intake with "deferred" rows,
the dirty-flag retry contract including a corrupt ledger, the count and time
bounds with intake reached while events remain queued, and /api/state serving
the same values for both budget branches without the full breakdown.
2026-09-25 03:21:41 +03:00
Ouroboros
aef127346b fix(benchmarks): retain official evaluator evidence independently of execution 2026-09-25 03:21:08 +03:00