Commit graph

2945 commits

Author SHA1 Message Date
Vyctor H. Brzezowski
06431c33f0
refactor(ui): model Workboard preview sessions consistently (#144747)
* fix(ui): preserve mock session lifecycle through send and abort

* fix(ui): retain mock cancellations and isolate Workboard fixtures

* fix(ui): reconcile mock session previews after cancellation

* fix(workboard): restore dashboard progress in mock

* refactor(workboard): isolate mock progress fixtures

* fix(ui): retain mock session failure diagnostics

* style(ui): format session diagnostic fixtures

* fix(ui): preserve newer mock session outcomes

* fix(ci): provision ripgrep for runtime projection tests
2026-09-13 17:11:33 -03:00
Peter Steinberger
0f8cf0f6c2
ci: record resources for slow runtime test commands (#147254) 2026-09-13 12:12:34 -07:00
RoboClaw
e0374f71a5
docs: require UI screenshots in chat and PR (#147263)
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-13 11:41:44 -07:00
Vincent Koc
07b96d9b41
fix(ci): retain QA profile diagnostics on failure (#143968)
* fix(ci): retain QA profile diagnostics on failure

* fix(ci): retain directly extracted QA shard diagnostics

Punchcard-Session: clear-meadow-meadow-4g

* fix(ci): use set lookup for QA diagnostic files

Punchcard-Session: clear-meadow-meadow-4g
2026-09-14 02:23:11 +08:00
Vincent Koc
6670690552
fix(ci): check publication registries before FRV fanout (#147184)
* fix(ci): check publication registries before FRV fanout

* test(ci): align FRV registry admission workflow fixtures
2026-09-14 02:07:41 +08:00
Peter Steinberger
149648a3d0
fix: publish Linux apps with stable releases (#147113)
* fix: publish Linux apps with stable releases

* test: include Linux updater in release workflow routing
2026-09-13 09:32:07 -07:00
RoboClaw
8baed117f6
fix(ci): preserve frozen repo E2E source identities (#146955)
Pass the already-validated selected source, trusted tooling SHA, and selected workspace through the reusable repo E2E workflow. Keep preflight validation and all scenario gates intact.

Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-13 08:59:04 -07:00
Vincent Koc
7bf4e2e088
fix(ci): reject invalid release sources before FRV fanout (#147049)
* fix(ci): reject invalid release sources before FRV fanout

* fix(ci): preserve supported FRV source admission routes
2026-09-13 22:10:51 +08:00
Peter Steinberger
1772bf1e8d
fix: allow versioned Linux release requests to publish (#146878)
Some checks are pending
Native App Locale Refresh / Refresh native it (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ja-JP (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ko (push) Blocked by required conditions
Native App Locale Refresh / Refresh native nl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pt-BR (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ru (push) Blocked by required conditions
Native App Locale Refresh / Refresh native sv (push) Blocked by required conditions
Native App Locale Refresh / Refresh native th (push) Blocked by required conditions
Native App Locale Refresh / Refresh native tr (push) Blocked by required conditions
Native App Locale Refresh / Refresh native uk (push) Blocked by required conditions
Native App Locale Refresh / Refresh native vi (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-CN (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-TW (push) Blocked by required conditions
Native App Locale Refresh / Commit native locale refresh (push) Blocked by required conditions
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Waiting to run
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Blocked by required conditions
Plugin NPM Release / preview_plugins_npm (push) Waiting to run
Plugin NPM Release / Validate release publish approval (push) Blocked by required conditions
Plugin NPM Release / preview_plugin_pack (push) Blocked by required conditions
Plugin NPM Release / Preflight plugin npm package () (push) Blocked by required conditions
Plugin NPM Release / Seal prepared plugin npm release (push) Blocked by required conditions
Plugin NPM Release / Trusted publisher OIDC exchange (push) Blocked by required conditions
Plugin NPM Release / publish_plugins_npm (push) Blocked by required conditions
Plugin NPM Release / verify_plugins_npm (push) Blocked by required conditions
Vitest Cache Warm / warm (linux) (push) Waiting to run
Vitest Cache Warm / warm (macos) (push) Waiting to run
Workflow Sanity / no-tabs (push) Waiting to run
Workflow Sanity / actionlint (push) Waiting to run
Workflow Sanity / generated-doc-baselines (push) Waiting to run
2026-09-13 02:11:42 -07:00
Peter Steinberger
055250e760
chore: cover custom-plugin sibling imports in full release validation (#146636)
* test: cover custom-plugin sibling imports in release validation

* fix: complete sibling scenario release harness integration
2026-09-12 19:11:00 -07:00
Peter Steinberger
db6b1ae5bc
refactor: share release artifact restore workflow definitions (#146528) 2026-09-12 18:46:32 -07:00
Peter Steinberger
b137b10eab
ci: skip duplicate PR corpus with complete Node coverage (#146465) 2026-09-12 15:10:38 -07:00
Peter Steinberger
5278a8f2ed
feat: share selected sessions read-only with a paired team Gateway (#136253)
* feat(node-host): advertise an explicit node command allowlist

Persist exact node command selection and restrict ancillary publication and hosting. Preserve the unchanged assertion baseline under the work-order stop rule; check:changed requests removing the obsolete runtime.ts count (2 to 0).

* feat(plugin-sdk): session transcript catalog reader

Expose bounded read-only native display pages and portable attribution through the existing runtime subpath. Keep pagination scoped to the original active transcript branch and allow an explicit bounded native cursor length.

* feat(session-share): read-only OpenClaw session catalog across paired gateways

Publish explicitly selected native session groups through two paired node commands. Validate the closed wire contract, reject remote profile claims, and keep receiver identity binding opt-in and display-only.

* fix(gateway): show published session catalogs to view-scoped roles

Let publication consent satisfy catalog read visibility for roles allowed to view others, while owner-only and unprofiled callers stay hidden. Preserve published attribution without accepting a remote local-session adoption claim. Regression tests reproduce four pre-fix failures; final validation stopped at the work-order baseline gate.

* docs: session sharing across gateways

Document sessions-only node setup, explicit publication groups, receiver attribution, view-scoped catalog access, and read-only limits. Add the bundled plugin inventory and generated reference entry. Live proof runbook remains outside the repository; build and rig execution are blocked by the work-order baseline restriction.

* fix(session-share): preserve source storage and paired reconnects

Respect configured stores through listing, paging, and revocation. Keep cold listings available and bound raw transcript reads. Prefer the established paired node credential on service restart, suppress unrelated host metrics, and refresh the approved plugin configuration docs.

* refactor(gateway): separate authorized catalog reads

Keep the catalog dispatcher within its owned scope and preserve post-read role checks and sender projection. Align the rebased tests with their shared setup and imports.
2026-09-12 13:35:17 -07:00
Peter Steinberger
d13f07b1c2
chore(deps): advance cooled dependencies and major upgrades (#146258)
* chore(deps): advance cooled dependencies and major upgrades

* test(logging): migrate failed-sink regression to tslog 5

* test: retain dependency upgrade coverage within lint limits

* fix(deps): preserve compiler launches, Matrix sync and chat metadata

Keep copied script harnesses independent of declaration modules and preserve
Windows executable prefixes after admission. Audit the Matrix sync guard for
42.3, align CI toolchain/cache pins, and refresh session facts after accepted
model-catalog invalidation without relying on picker timing.

* fix(ui): preserve scoped session reconciliation after catalog refresh
2026-09-12 13:28:12 -07:00
RoboClaw
d23730ee7c
docs: make PR descriptions plain-language first (#146253)
Co-authored-by: hannesrudolph <49103247+hannesrudolph@users.noreply.github.com>
2026-09-12 13:29:17 -06:00
Peter Steinberger
77292356e9
fix(ui): avoid redundant child-session refreshes in chat (#146122)
* fix(ui): avoid redundant child-session refreshes in chat

* fix(ui): preserve scoped chat refresh ownership

* test(ui): keep scoped session fixtures authoritative

Keep canonical mock transitions and child query membership aligned across descriptor, history, and list reads. Preserve real Gateway stale Stop coverage by delaying only the exact terminal descriptor until post-Stop history recovery. Rebalance existing Gateway test-type graphs after main exceeded the unchanged root budget.

* ci: match real Gateway job budgets to hosted runners

Give the existing hosted real-Gateway route a 40-minute job budget while retaining 20 minutes on Blacksmith. The hosted diagnostic run reached its job deadline after 7m42s of preparation and a progressing full suite. Preserve every per-test deadline, test inventory, worker, runner and concurrency setting; check budget/routing agreement across 96 contexts.
2026-09-12 11:40:00 -07:00
Ayaan Zaidi
7753ffb475
fix(ui): publish initial model catalog on connect (#145927)
## What Problem This Solves

The first web model picker open could wait seven seconds for a catalog reply. A late initial snapshot could also restore an old saved account after the user changed it.

## Why This Change Was Made

Publish the existing catalog during connection setup through the authenticated request dispatcher. Initial and ordinary replies use the same cache publication checks, so delayed work cannot replace a newer session selection.

## User Impact

Published choices appear on first open and reconnect while discovery continues. Warm reopen stays fast, without an extra catalog request from opening the picker.

## Evidence

With catalog replies held for seven seconds, the original picker had no usable choices until the reply. The corrected New Session and active-chat pickers showed rows within 100 ms; warm reopen took 8–12 ms. Neither open added a request.

Real browser/Gateway tests cover reconnect and a saved-account change before snapshot completion. Sixteen valid configuration fixtures start; an invalid legacy fixture rejects as expected. Previous-version source builds connect in both directions; installed-release upgrades were not tested. Full consumer analysis: `consumers.md`.

## Compatibility

The new connect field is sent only after server advertisement. Clients without opt-in receive no snapshot. No protocol-version bump, configuration key, database change, or migration.

## Consumers

- New Session picker: uses published rows before its ordinary catalog reply.
- Active-chat picker: receives the saved session's model/account projection and retains newer selections.
- Mounted browser shell: preserves initial publication across reconnect.

## Invalidation

Connection, identity, configuration, session metadata, session mutations, and explicit refresh retire stale reads. Expiry and eviction retain publication ordering; delayed results cannot refill invalidated state.

## Contention

Snapshot publication adds no lock, transaction, or await. The existing dispatcher owns catalog reads; the common cache owner admits results synchronously.

## Tests

Browser picker tests exercise fresh, saved-session, refresh, and reconnect flows. Gateway tests enter connect, `models.list`, and `sessions.patch`. Handshake, cache, metadata, and sibling checks pass. The shared test partition guard passes with the server grouping from #146212. Full automated validation is still required.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-12 23:39:39 +05:30
RoboClaw
2a1bc8be3d
feat(ui): watch live desktops in Picture-in-Picture (#143836)
Add native Document Picture-in-Picture for live desktop observation without opening another connection or changing input ownership. Keep the mirror view-only, capability-gated, and tied to viewer teardown, with stale asynchronous requests safely retired.

Preserve current desktop sizing and lock notices. Cover lifecycle races, native PiP and focus-loss liveness; publish only synthetic before/after evidence through CI. Correct the real-gateway toolbar expectation for the new PiP control.

Proof: https://github.com/openclaw/openclaw/actions/runs/34703111201/artifacts/10301576157

Native Linux Chromium minimization/stacking was verified in isolation. Native macOS/Windows and genuine Firefox background scheduling remain unverified; Linux proof is not generalized.

Closes #143829

Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-12 09:16:09 -07:00
Peter Steinberger
2ea6c54476
fix(ci): balance slow runtime test jobs across existing runners (#145956)
* perf(ci): balance runtime test groups across existing jobs

* fix(ci): invalidate timing publication when its decoder changes
2026-09-12 08:17:28 -07:00
RoboClaw
e51bc4de5f
fix: restore explicit Sol coverage in full release plans (#145974)
Restore the original native Ultra model selection and stable/full Codex Docker cell without changing production defaults, runner budgets, credentials, or proof gates.

Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
2026-09-12 06:25:59 -07:00
Ayaan Zaidi
2e4391ac95
fix(auth): recover accounts after subscription quota resets (#145853)
Related: #123009

## What Problem This Solves

A subscription can remain blocked for days after its provider usage view reports restored capacity. The next ordinary chat, scheduled run, or inherited agent turn keeps rejecting the same valid account. The stored provider reset deadline remains unchanged.

## Why This Change Was Made

Reconcile stored subscription blocks through the existing auth usage owner before runtime admission. After five minutes, the next demand waits for a fresh usage observation. Both persisted quota sources can recover. Additional quota buckets, workspace limits, spending controls, malformed responses, and account failures cannot falsely clear the block.

Concurrent demands share one probe per physical credential owner. The apply transaction verifies the current credential and unchanged failure generation. Each waiting caller then reads its current persisted credential and usage. Direct side questions use the same owner before preparing auth. The isolated-completion correction moves host auth through common simple-completion preparation, which awaits quota reconciliation before route selection and the immutable auth plan. The native-auth branch retains the same quota owner. Caller authority is checked again after the wait.

A fresh available usage response received during another model's ordinary failure uses that same quota transition immediately. It clears the old quota restriction before accounting for the new ordinary failure. Delayed usage responses cannot overwrite a newer quota observation, credential replacement, or successful use.

Successful use records its health observation in the existing `lastProbeAt` field. This also protects inherited accounts, whose successful turns intentionally preserve account-selection history.

Generic rate-limit accounting no longer adds a second retry cooldown after the provider quota owner has recorded that failure. Existing rows with both restrictions can reach the quota check; a fresh result retires only a rate-limit cooldown whose model scope is covered by the quota block. Authentication, disabled-account, and unrelated model cooldowns remain effective.

Operational SQLite write failures during quota reconciliation retain the stored block and allow a healthy configured fallback. The shared pending probe handles only native READONLY, IOERR, and FULL errors. The normal store writer records the error; each caller must still reload current credentials and state before admission. Constraint and other unexpected errors remain fatal.

Token renewal remains with the existing credential owner. No configuration option, schema, migration, provider plugin, or public SDK contract changes.

The current combined production diff is net-negative: 449 added lines and 467 removed lines across eight production files, a reduction of 18 lines. Tests and support are counted separately. Current signed head is `625c2c137262c475e1bec74e2ede9287672116f0`, tree `b4de089eecac78ff5468bd3dbd63e91f9768ef2f`. Product source is byte-identical to the successful correction proof at `b6ec779939db15a7ffd56aba4bb539d1ea9a2f7d`; the later test/CI follow-up passed its own affected checks. Signature and source bindings are verified; fresh exact-head integration checks remain required.

This changes the background-only admission choice introduced in #108683 (`a5d1d9e473babae7ebcb4a48c958ec5158bc9e2e`): the current demand now waits for the bounded usage check and can use restored capacity immediately. @steipete, for awareness.

## User Impact

After the five-minute recheck interval, the next ordinary turn can use the same account when upstream capacity is restored. Users do not need to restart, delete state, or sign in again to clear a stale quota block.

## Consumers

- `src/agents/auth-profiles/types.ts`, `profile-usage-stats.ts`, `state.ts`, and `sqlite.ts` define, normalize, merge, and persist `blockedUntil`, `blockedReason`, `blockedSource`, `blockedModel`, `blockedScope`, and `lastProbeAt`; their shapes are unchanged. The `lastProbeAt` comment now documents a quota probe or successful provider use.
- `auth-profiles/usage.ts` owns failure accounting, probe claims, and guarded quota reconciliation. `usage-state.ts` clears expired state and supplies synchronous admission/display predicates; its existing model-scoped block predicate now also serves accounting and reconciliation. The existing `profiles.ts::markAuthProfileSuccess` transaction now clears health and stamps `lastProbeAt` together. Its single-use `updateSuccessfulUsageStatsEntry` helper is removed. `upsert-with-lock.ts` retains explicit failure reset when credential replacement requests it.
- `embedded-agent-runner/run/prompt-failure.ts`, `auth-controller.ts`, and `failover-retry-controller.ts` route generic failure accounting after the native quota writer. `cli-runner/cli-run-settlement.ts` reaches the same owner through `cliRunSettlementDeps`; `cli-runner.ts` retains the existing dependency alias and test reset. The usage owner no longer counts the same covered rate-limit failure twice or installs a competing cooldown. Saved profiles retain ordinary rate-limit backoff when no subscription block covers the failure. Inline-key accounting accepts only auth, permanent-auth, and billing failures; those policies remain unchanged.
- `auth-profiles/store.ts`, `store-runtime.ts`, and `runtime-snapshot-owner.ts` own effective credential inheritance, transactional updates, and snapshot invalidation. Reconciliation keys work by physical owner, then reloads each caller's effective store after waiting. The shared pending probe now contains only operational SQLite write errors; the generic store contract is unchanged. `src/infra/sqlite-error-diagnostics.ts::sqlitePrimaryResultCode` is reused unchanged to identify native error codes. The catch applies equally to model fallback, side questions, and isolated completion; their mandatory effective reload remains outside it.
- `src/state/user-model-accounts.ts` normalizes and persists personal-account `usageStats` through `coerceProfileUsageStats`. `auth-profiles/personal-profiles.ts` materializes only the selected account and applies the existing credential/usage updater inside that account's transaction. This retains exact-account ownership and the existing state shape; it does not copy personal credentials into shared or agent stores.
- `model-fallback-runner.ts` is the existing production caller of `maybeReprobeWhamBlockedProfiles`; it now awaits the operation for native and ordinary harnesses. `auto-reply/reply/agent-runner-fallback-candidate.ts`, `embedded-agent-runner/run-entry.ts`, and `cron/isolated-agent/run-executor.ts` route chat, inherited agent, and scheduled work through that boundary. Compaction reaches it through `embedded-agent-runner/compact.ts`.
- `btw.ts` invokes the quota owner before direct auth plans. Host V2 and legacy isolated completion now enter `simple-completion-runtime.ts::prepareSimpleCompletionModelCore`, which awaits reconciliation before route selection and auth planning. `isolated-completion.ts` carries its admitted authority check into that wait and revalidates afterward; the native-auth branch retains its earlier shared-owner call. `runtime-plan/prepare-auth.ts`, `runtime-plan/resolve-auth.ts`, `embedded-agent-runner/run/auth-plan.ts`, `runtime-preparation.ts`, and `auth-controller.ts` retain final admission checks. Native bound-session auth remains owned by its harness.
- `auth-profiles/order.ts` and `session-override.ts` consume unchanged failure state for ordering and preferences. Hard user selections remain scoped; automatic preferences still permit reconciliation of the full eligible order.
- `model-fallback-cooldown.ts` consumes reset metadata for retry scheduling. It does not mutate quota facts. `model-catalog-decisions.ts`, `model-auth-availability.ts`, and `model-catalog-auth-labels.ts` derive availability and display state without network requests.
- `model-fallback-attempt.ts` reloads persisted auth state before computing per-candidate cooldown expiry for the final failure summary. It preserves the selected personal-account scope and uses the same synchronous expiry reader.
- `src/commands/models/list.auth-overview.ts`, `list.status-command.ts`, `auth-list.ts`, and `doctor-auth.ts`, plus `extensions/codex/src/command-account.ts` and `extensions/discord/src/monitor/auto-presence.ts`, render the same state. `src/gateway/server-methods/models-auth-status.ts` exposes last-use metadata when profile identity is included. `auth-profiles/state-observation.ts` logs cooldown transitions and failure counts without owning admission. `extensions/codex/src/app-server/usage-limit-error.ts` remains the native block writer, called from turn start/finalization. Native rate-limit cache data does not authorize recovery.
- `src/agents/auth-profiles.ts`, `auth-profiles.runtime.ts`, and `src/plugin-sdk/agent-runtime.ts` retain their existing exports. The runtime facade exposes the awaited probe; the general facade exposes failure and success accounting; the SDK exposes native block writing and synchronous state helpers. The new direct-admission operation and shared model-scope predicates remain internal. Lexical field matches include unrelated channel-health and UI mutation-block fields. The actual provider-status browser consumers are named below. Existing database and compatibility documentation retains the same persisted contract. Typed-reference coverage is limited to six loaded projects; 134 other discovered configurations were not selected, and native code, dynamic keys, and external consumers are not excluded by that census.
- Private removed helpers/types/constants have no remaining production callers: `claimWhamHalfOpenReprobe`, `WhamBlockGeneration`, `WhamUsageResponse`, `resolveWhamCooldownClassification`, and `FAILURE_REASON_ORDER`. Their ownership is absorbed by the existing probe operation, validated response schema, typed classification, and priority-ordered selection.
- The private probe result's generic `reason` strings had only one classification consumer in `usage.ts`. The probe now carries the existing typed auth classification directly; diagnostic and persisted classification values remain unchanged. Its private `blockedSource` property was always WHAM and is now set directly by that writer; the public persisted field is unchanged. Retired non-auth probe strings have no remaining consumers.
- `resolveWhamCanonicalCooldownReason` had one private caller in `applyWhamCooldownResult`; its unchanged classification mapping now lives there. `resolveAuthCooldownConfig` and its five private constants had only two accounting callers, `markAuthProfileFailure` and `markInlineProviderApiKeyFailure`. Their private `cfgResolved` fields are removed from `computeNextProfileUsageStats` and `resolveDisabledFailureBackoffMs`. One static `AUTH_COOLDOWN_CONFIG` retains every duration. The private `DISABLED_FAILURE_BACKOFF_POLICIES` records now contain numeric `baseMs` and `maxMs` values, selected by failure reason and passed to the unchanged exponential-backoff calculation; no configurable or public policy surface is removed.
- `lastUsed` remains selection history for profile ordering, account selection, and Gateway metadata; it is not part of quota generation matching. Successful use now stamps the existing health field `lastProbeAt` inside its owner transaction, including inherited use that preserves `lastUsed` and `lastGood`. Embedded terminal resolution reaches this owner through `markEmbeddedRunAuthProfileSuccess`; CLI settlement awaits it through `cliRunSettlementDeps`, also aliased by `cliRunnerDeps` for existing test reset. `markAuthProfileFailure` reads the current effective persisted store before probing; admission and failure paths compare health generation and credential identity before quota writes.


The isolated text consumers below are in scope for their common auth-admission behavior. Session assessment, summaries, progress narration, labels, and plugin outputs retain their existing output owners. A registered task pass does not become a separate runtime pass for every sibling output.

| Production/pipeline path and source | Role, scope, and disposition |
| --- | --- |
| `src/agents/simple-completion-runtime.ts::prepareSimpleCompletionModelCore` | Host simple/isolated preparation now waits for the same quota owner before route selection and immutable auth planning. The selected store/model/profile and caller authority remain scoped. Registered restored-capacity failure is retained on the unchanged base; the tested correction passes restored recovery, due-exhausted refusal, and removed-profile controls. |
| `src/gateway/server-plugin-subagent-runtime.ts:274` | Authorized plugin-host isolated execution calls the common owner, passes `assertCurrent`, and returns text after authority/signal checks. In scope for shared admission. The registered task proof reaches the plugin isolated route; no standalone subagent-control campaign is implied. |
| `src/plugins/runtime/runtime-llm-isolated.ts:152` | Plugin completion adapter forwards selected account, model, timeout, and cancellation to the same isolated owner. In scope for shared admission and directly exercised by the registered isolated-task proof; its error translation/cleanup remains unchanged. |
| `src/gateway/session-observer-model.ts:266` | Default completion delegates to isolated execution for operator session assessment. In scope for shared admission; observer parsing, acceptance, and digest publication are unchanged and not separately claimed as quota runtime proof. |
| `src/transcripts/summary-model.ts:127` | Configured summary attempts preserve selected profile and deadline, then parse the isolated completion into meeting notes. In scope for shared admission; existing summary fallback/format contracts stay unchanged. No transcript-summary-specific runtime pass is claimed. |
| `src/transcripts/summary-model.runtime.ts:1` | Lazy re-export of the common isolated operation and selection helper. Same in-scope summary route; no second admission or quota owner. |
| `src/auto-reply/reply/progress-narrator-model.ts:84` | Prepared utility-model narration calls isolated completion under its cancellation/timeout and returns text or its existing failure result. In scope for shared admission; narration scheduling and output handling remain unchanged. |
| `src/auto-reply/reply/conversation-label-generator.ts:130` | Label attempts pass exact selected/preferred account and compatible runtime to isolated completion, then revalidate authority before accepting text. In scope for shared admission; label formatting and candidate policy remain unchanged. |
| `src/agents/embedded-agent-runner/compaction-runtime-preparation.ts:173`, `:212`, `:229` | Fresh compaction auth plans consume the effective store; a reusable plan retains its already selected identity. Ordinary compaction runs through `compact.ts:507` → model fallback before direct preparation. `direct-compaction-preparation.ts:172` and `compact.queued.ts:505` reuse this preparation. In scope for that existing admission/plan lifecycle; no independent probe or recovery state is added to queued/reused-plan preparation, and no separate queued-compaction quota pass is claimed. |
| `extensions/codex/src/command-rpc.ts:118` | Native control subscriptions select the same auth partition as the admitted session and enforce its generation/authority. In scope for selected-account and refresh consistency. A control subscription is not a new model inference; this path does not own an additional quota probe. Its control-only recovery is source-inspected, not a new runtime claim. |
| `extensions/codex/src/app-server/rate-limits.ts:231`, `:239`, `:283` | Native limit payloads supply the reset-time writer input and account/usage projections. In scope as upstream facts and display data. Cached “available” display is not permission to clear persisted quota; fresh recovery remains with the quota owner. |
| `extensions/codex/src/app-server/rate-limit-cache.ts:18`, `:39`, `:66`, `:75` | Per-physical-client reads and sparse notifications update revisioned cached limits. Readers enforce max age. Existing cache/display/native-failure contracts remain; it does not become a second quota recovery authority. |
| `src/secrets/runtime.ts:650` | Active provider-auth refresh prepares candidate snapshots and activates only while config/auth revisions still match. Both the retained base and tested correction observe successful ordinary chat, command-line status, and actual provider-header/profile Ready status after saved clearing and refresh. The correction run also passes exact inference bearer/account checks. |
| `src/secrets/runtime-state.ts:261`, `:307` | Activation merges live bookkeeping for matching owners and recomposes on owner transitions, preserving order, lastGood, and usageStats. Saved-clear/refresh behavior passes ordinary-turn, command-line, and browser checks on the recorded correction sources as well as the base. No additional state owner is introduced; the tested sources match the signed head. |
| `src/gateway/model-auth-refresh.ts:13` | Post-mutation refresh reloads shared ownership, clears status usage cache, refreshes provider auth, records mutation, and prepares runtime state. It remains distinct from the explicit `models.authStatus({refresh:true})` path used for the C2 manual saved-clear observation. No refresh acknowledgement alone is treated as final-consumer proof. |
| `src/gateway/server-methods/models-auth-refresh.ts:10` | Registered `models.authRefresh` validates scope and awaits the post-mutation refresh before returning its acknowledgement. In scope; acknowledgement is not final evidence of ordinary-turn or UI convergence. |
| `ui/src/lib/model-auth.ts:70` | Browser `models.authStatus` reads share only current flights. Explicit refresh retires pending sharing; settled replies are not cached here. Both correction browser cases show provider-header and profile Ready status, with exact inference account checks and inspected captures. |
| `ui/src/pages/model-providers/load.ts:82`, `:95` | The page waits for auth refresh before configured-catalog refresh, then joins auth/catalog/config results. In scope for the final status pipeline; both real-browser cases pass on the recorded correction sources. |
| `ui/src/pages/model-providers/data.ts:267` | Derives provider cards, selected profile identities/order, auth and usage data from loaded responses. It does not clear health. In scope for final displayed provider facts; both real-browser cases pass on the recorded correction sources; final captures were inspected. |
| `ui/src/pages/model-providers/model-providers-page.ts:639`, `:654` | Supplies loaded cards/auth status to the view and routes explicit refresh through the page owner. In scope for the visible provider page; both real-browser cases pass on the recorded correction sources; final captures were inspected. |
| `ui/src/pages/model-providers/view.ts:545`, `:600` | Renders provider rows and model readiness from those facts. In scope for final visible status, not merely a responsive request; both real-browser cases pass on the recorded correction sources; final captures were inspected. |

The common simple-preparation correction also reaches these existing consumers:

| Consumer | Disposition and proof limit |
| --- | --- |
| `src/plugins/runtime/runtime-llm.runtime.ts:588` | Direct-provider plugin completion now awaits due quota facts through model acquisition; mode/profile checks and resource drainage remain. The registered isolated task does not execute this direct-provider branch. |
| `src/cli/capability-cli/model.ts:204` | Local `infer model run` acquires the same prepared model before direct execution. This is a source-covered inference consumer; C2's `models status --json` is not its runtime proof. |
| `src/plugin-sdk/simple-completion-runtime.ts` | The public ForAgent facade delegates to the repaired acquisition owner and preserves the existing resource-host lifetime. Its parameter/return schema remains unchanged; no specific SDK consumer runtime pass is claimed. |
| `src/tts/tts-core.ts:203` | Default long-speech text-summary preparation gains the same due-quota recovery. Its request timer still begins after preparation; the bounded probe adds preparation latency. Custom injected preparation retains its existing contract. |
| `src/tts/tts-payload.ts:216` | Successful summarization feeds speech text; preparation failure retains existing truncation behavior. No synthesis, audio, or channel-result pass is claimed by the isolated task test. |
| `src/plugin-sdk/speech-core.ts` | The exported summary dependency keeps its unchanged caller-owned preparation type. The lifecycle callback stays a separate private argument; custom injected implementations are not silently changed. |

Personal accounts remain in scope for exact selected-account recovery and ownership through `src/state/user-model-accounts.ts` and `src/agents/auth-profiles/personal-profiles.ts`. Their source/owner-test evidence does not establish a personal-login end-to-end quota pass. The shared pending-key prefix never moves personal credentials into shared or agent stores.

The exact retained typed test/support inventory follows. “Unchanged” compares the authored base to the published pre-correction head; only `usage.test.ts` among these 40 paths changed in that comparison. These rows name source contracts and do not request or claim 40 new test runs.

| Exact test/support path and inspected source | Contract and disposition |
| --- | --- |
| `src/agents/auth-profiles.cooldown-auto-expiry.test.ts:28` | Unchanged; ordering expires old cooldowns, keeps active siblings, and preserves rate-limit history until success. |
| `src/agents/auth-profiles.getsoonestcooldownexpiry.test.ts:164` | Unchanged; retry-time reader includes profile-wide blocks while excluding unrelated model-scoped cooldowns. |
| `src/agents/auth-profiles.markauthprofilefailure.test.ts:87` | Unchanged; failure persistence must not overwrite a fresher stored credential. Ordinary billing and timeout accounting remains a sibling contract. |
| `src/agents/auth-profiles.resolve-auth-profile-order.does-not-prioritize-lastgood-round-robin-ordering.test.ts:182` | Unchanged; lastUsed determines round-robin order; lastGood must not override it. |
| `src/agents/auth-profiles.sqlite-store.test.ts:128` | Unchanged; credential/order/lastGood/lastUsed round-trip through the existing SQLite store. |
| `src/agents/auth-profiles/catalog-order.test.ts:67` | Unchanged; catalog credential selection retains configured order and model-scoped cooldown behavior. |
| `src/agents/auth-profiles/oauth.concurrent-agents.test.ts:169` | Unchanged; one durable refresh owner removes copied credentials while preserving each agent's selection history. |
| `src/agents/auth-profiles/oauth.fallback-to-main-agent.test.ts:362` | Unchanged; a hard selected profile retains refresh failure, while unlocked selection can use a healthy sibling and settle success. |
| `src/agents/auth-profiles/order.test.ts:395` | Unchanged; valid profile ordering and lastUsed inputs stay distinct from quota health timestamps. |
| `src/agents/auth-profiles/persisted-boundary.test.ts:51` | Unchanged; persisted reason/classification normalization is shared by runtime-state reading and writing. |
| `src/agents/auth-profiles/portability.test.ts:40` | Unchanged; portable credential copies omit usageStats and lastGood; quota recovery does not copy owner health. |
| `src/agents/auth-profiles/profiles.test.ts:1488` | Unchanged test, affected success owner; successful use clears classification/backoff through the existing profile transaction. |
| `src/agents/auth-profiles/runtime-snapshots.test.ts:192` | Unchanged; usage-only changes do not masquerade as credential-ownership notifications. |
| `src/agents/auth-profiles/session-override.rotation.test.ts:92` | Unchanged; automatic selection retries the preferred profile after cooldown expiry. |
| `src/agents/auth-profiles/session-override.test.ts:640` | Unchanged; provider/model blocks and healthy siblings participate in explicit versus automatic session selection. |
| `src/agents/auth-profiles/state-observation.test.ts:16` | Unchanged; diagnostic formatting sanitizes fields without becoming a state writer or admission rule. |
| `src/agents/auth-profiles/store-owner-publication.test.ts:289` | Unchanged; committed credential/state facts remain recorded when migration refuses derived publication. |
| `src/agents/auth-profiles/store-state-owner.test.ts:42` | Unchanged; shared/local/mixed bookkeeping ownership survives shared-root activation. |
| `src/agents/auth-profiles/upsert-with-lock.sqlite.test.ts:493` | Unchanged; explicit credential replacement resets target health while preserving selection history and sibling state. |
| `src/agents/auth-profiles/usage.inherited-owner.test.ts:309` | Unchanged; inherited success clears health without moving selection ownership. Personal transaction/isolation cases at :103/:157/:233 remain source-scoped evidence. |
| `src/agents/auth-profiles/usage.test.ts:1322` | Changed earlier in this PR; owner-level persistence mocks and four Promise consumers adapted. Three exact case purposes and registered counterparts are listed in Tests below; no assertion deleted in those adaptations. |
| `src/agents/cli-runner.before-agent-reply-cron.test.ts:294` | Unchanged; recovered command execution records success and clears stale health, while pre-execution denial must not settle health. |
| `src/agents/embedded-agent-runner.run-embedded-agent.auth-profile-rotation.e2e.test.ts:683` | Unchanged; read-only runs do not persist bookkeeping; surrounding cases retain auth rotation and lastUsed behavior. |
| `src/agents/embedded-agent-runner.run-embedded-agent.fault-sequences.e2e.test.ts:502` | Unchanged; real runner fault sequencing preserves provider/profile fallback and eventual success bookkeeping. |
| `src/agents/embedded-agent-runner/run.overflow-compaction.misc-owners.test.ts:28` | Unchanged; the terminal success wrapper keeps its asynchronous bookkeeping contract and does not wait on the owner. |
| `src/agents/embedded-agent-runner/run/assistant-failure.test.ts:261` | Unchanged; actual inference-storage failure stays terminal without credential rotation/replay. Optional quota-recovery I/O handling is a different, earlier boundary. |
| `src/agents/embedded-agent-runner/run/auth-controller.test.ts:597` | Unchanged; existing disabled health prevents auth initialization and preserves the corresponding failure classification. |
| `src/agents/embedded-agent-runner/run/terminal-resolution.test.ts:185` | Unchanged; terminal success reports the selected profile privately, including read-only command-maintenance flows. |
| `src/agents/model-auth-availability.reasons.test.ts:181` | Unchanged; availability/next retry excludes invalid or permanently rejected profiles from viable cooling candidates. |
| `src/agents/model-auth-availability.session-pin.test.ts:62` | Unchanged; personal/shared pin projections match prepared-runtime candidate selection, including cooldown and missing credentials. |
| `src/agents/model-auth-availability.test.ts:326` | Unchanged; a cooling automatic tier retains its known physical auth route while reporting unavailable. |
| `src/agents/model-auth.profiles.test.ts:1327` | Unchanged; a configured inline key's billing disable remains enforced; subscription probing does not replace inline-key policy. |
| `src/agents/model-catalog-auth-labels.test.ts:112` | Unchanged; captured labels/order change when the captured cooldown expires; labels do not authorize fresh quota recovery. |
| `src/agents/model-fallback.personal-cooldown.test.ts:45` | Unchanged; fallback summaries reload the exact selected personal account's newly persisted cooldown. |
| `src/agents/model-fallback.probe.test.ts:424` | Unchanged; single-provider long-block retry scheduling and fallback preference remain separate from quota-observation persistence. |
| `src/agents/model-fallback.run-embedded.e2e.test-support.ts:77` | Unchanged support; writes synthetic credential/order/usage fixtures for the combined fallback/runner suite; no production state owner. |
| `src/agents/model-fallback.run-embedded.e2e.test.ts:674` | Unchanged; combined fallback and runner outcomes retain correct provider attribution, clean failure state, and successful sibling lastUsed. |
| `src/agents/model-fallback.test.ts:841` | Unchanged; exhausted-chain summaries preserve selected-profile identity; surrounding cases retain fallback ordering and shared health fixtures. |
| `src/agents/runtime-plan/prepare-auth.test.ts:459` | Unchanged; immutable auth planning fails closed for all-cooldown profile orders. |
| `src/agents/runtime-plan/resolve-auth.test.ts:462` | Unchanged; prepared auth resolution rechecks model-scoped cooldown before selecting a stored candidate. |



The completed correction census adds 32 paths beyond the original 80; all original paths remain. Sixteen added paths are already dispositioned in the isolated, direct, SDK, speech, and registered-test entries above. The remaining sixteen are named here. Approval, worker placement, and output-specific gates retain their own authority; the isolated-task pass is not their runtime proof.

| Additional census path | Scope and disposition |
| --- | --- |
| `src/agents/embedded-agent-runner/run.overflow-compaction.test.ts` | Typed acquisition mock exercises compaction composition, locked model/auth/fallback facts, recovery outcome, usage, and transcript ownership. Retained owner-level test; it does not execute quota admission. |
| `src/agents/exec-auto-review.stress.test.ts` | Injected acquisitions test malformed/concurrent outcomes, cancellation and timeout isolation, and one-shot decision binding. No real command approval or quota recovery is inferred. |
| `src/agents/exec-auto-reviewer.resources.test.ts` | Actual reviewer acquisition with a local provider checks held preparation/response/cancellation tails and disposal after drain. Retained lifetime sibling, not a subscription-reset case. |
| `src/agents/exec-auto-reviewer.test.ts` | Injected dependency tests retain conservative decision parsing, reviewer selection, unavailable/error fallback, timeouts, cancellation, and one-shot restrictions. They do not execute the quota reader. |
| `src/agents/exec-auto-reviewer.ts` | Configured command/widget reviewer acquisition gains due-quota recovery before model judgment. Ask/deny/timeout, low-risk allow-once restrictions, caller approval policy, and resource drainage remain unchanged. Available quota never supplies approval; actual permission decisions remain source-covered. |
| `src/agents/isolated-completion.resources.test.ts` | Actual local-provider host calls protect overlapping auth/response/cancel tails, distinct forks, generation retention, and disposal. No provider-created subscription block is part of this lifetime suite. |
| `src/agents/isolated-completion.test-support.ts` | Registers controlled model/harness/command/preparation mocks before importing the real isolated owner. Unit infrastructure accepts the trusted callback but is not a registered runtime entry point. |
| `src/agents/simple-completion-runtime.generation.test.ts` | Retains generation-bound model/auth/route consistency, exact selection, canonical utility/default selection, and disposal after canceled preparation. Constrains the new wait without replacing registered quota proof. |
| `src/agents/simple-completion-runtime.plugin-scope.test.ts` | Cold fixtures retain selected plugin-generation loading, acquired/borrowed ownership, prepared transport/credential identity, and local-service reload before inference. Source-covered lifecycle/transport sibling. |
| `src/agents/simple-completion-runtime.selected-model.test.ts` | Local-provider requests verify once-normalized model IDs through agent, raw, and worker-shaped execution. Protects canonical model scope; it is not a full placed-worker quota campaign. |
| `src/agents/simple-completion-scope.ts` | Resolver adapter binds model resolution to the caller's exact prepared generation/store pair. Canonical resolved model identity reaches reconciliation without a new ambient selection owner. |
| `src/agents/simple-completion.types.ts` | Prepared-model, agent-selection, and public ForAgent parameter/return aliases keep their shapes. The internal authority callback does not become a public selection option. |
| `src/gateway/worker-environments/inference-runtime.ts` | Approved model/session profile reaches shared preparation with exact credential binding. Catalog visibility, session identity, owner epoch, and pre-stream/per-event liveness still govern execution/publication. Preparation may finish after cancellation; no prompt-cancellation or full placed-worker runtime claim is made. |
| `src/plugins/contracts/tts-contract-suites.ts` | Reusable Vitest infrastructure despite its directory: injected speech-summary preparation checks summary text/metrics, selected override, transport, limits, and missing output. It is not a production quota owner or real recovery test. |
| `src/system-agent/approval-intent.ts` | The existing verified-route classifier may recover due quota only for its explicit bound credential. Route/fingerprint checks before and after inference, closed-list shortcuts, and failure-to-other behavior remain; proposal consent is never inferred from quota availability. Actual consent/persistence output remains source-covered. |
| `src/tts/runtime-api.ts` | The testApi.summarizeText alias points to the existing summary function. It adds no preparation or quota policy; synthesis and stream exports keep their own owners. |


The completed product-correction candidate and integration census contains 654 query/project pairs, 648 resolved queries, 5,783 reference rows, and 112 paths, with zero process failures. It retains all original 80 paths and adds the 32 named consumers. The base correction supplement contains 120 pairs, 117 resolved queries, 673 references, and 32 paths; its overlap with the old inventory differs because isolated production was already present while the new registered isolated test appears only in the candidate.

Six projects are loaded: `tsconfig.core.json`, `tsconfig.extensions.json`, `tsconfig.ui.json`, `test/tsconfig/tsconfig.core.test.agents-other.json`, `test/tsconfig/tsconfig.core.test.agents-root.json`, and `test/tsconfig/tsconfig.test.root.json`. The retained coverage records name 134 unselected configurations. Dynamic, native, browser, and other non-typed flows have the explicit source dispositions above; this is not whole-repository typed coverage.

| Unresolved query | Exact project-scope limits |
| --- | --- |
| Historical and still-current `src/plugin-sdk/agent-runtime.ts:133:3`, `markAuthProfileBlockedUntil` | Three gaps remain in `tsconfig.ui.json`, `test/tsconfig/tsconfig.core.test.agents-other.json`, and `test/tsconfig/tsconfig.core.test.agents-root.json`. The correction supplement does not erase them. |
| Newly queried `src/plugin-sdk/simple-completion-runtime.ts:14:14`, `prepareSimpleCompletionModelForAgent` | Three additional gaps occur in the same UI/agents-other/agents-root projects. Core, extensions, and root tests resolve the facade; its manual adapter tracing is recorded above. |

All six unresolved query/project pairs remain explicit coverage gaps. Neither an outside-project source nor zero candidate occurrences establishes zero consumers. The completed census closes the named inventory, not a formal review/merge approval.

The browser regression also has these test-pipeline consumers: `.github/workflows/ci.yml` now includes it once in the actual real-Gateway command; `test/vitest/vitest.ui-e2e.config.ts` keeps its canonical private-server membership sorted; `test/vitest/vitest.ui-e2e-prebuilt.config.ts` consumes that classification and leaves the new file outside its parallel allowlist; `test/vitest/vitest.e2e.config.ts` consumes the same real-Gateway list for exclusion from the generic E2E project. The existing workflow and partition guards below protect these contracts. Browser startup consumes the executable selected by `test/vitest/vitest.ui-e2e.global-setup.ts` through `controlUiE2eChromium`; `ui/src/test-helpers/control-ui-e2e.ts` remains the provisioning/path owner. The test no longer requests an unrelated default browser binary. No runner count, resource cap, timeout, retry, or product policy changes.

## Contention

The existing auth store transaction claims an admission probe and later applies its result. The network request runs between those transactions. Credential saves, other failure accounting, success updates, ordering, and shared snapshot publication use the same store owner. Success clears health and records `lastProbeAt` in that same transaction; the embedded success wrapper retains its existing asynchronous bookkeeping contract. The in-flight map coalesces admission probes by physical owner and profile within one process; failure-triggered usage probes remain separate requests and rely on the generation check when applying available or exhausted observations. Persisted claims throttle later probes across processes, but only callers in the same process share the pending result. Independent Gateway proof observed one held restored-usage request shared by main, a spawned child, and a registered cron run; all three subsequently succeeded. `health`, `cron.status`, and `models.list` completed while the request was held.

For personal accounts, the owner-directory resolver returns no agent directory. The in-flight key uses the shared path prefix plus the unique personal profile ID, so agents using the same personal account share its pending admission probe. Writes still route through the personal-account transaction, and each effective store materializes only the selected account. This source inspection does not add personal-login runtime proof.

## Invalidation

- Fresh usage or a bounded demand retry replaces or clears quota facts only through the quota owner.
- The provider's recorded reset time still expires through the existing synchronous state owner.
- Successful login, token renewal, credential replacement, account deletion, and a newly local inherited credential use the existing credential/store publication owners; waiting callers reread the current effective store.
- A newer failure generation or credential change prevents the old probe from applying.
- A committed successful use or quota probe updates `lastProbeAt`, which participates in generation matching even when no quota block existed before the request. Independent main and genuine-child overlaps verified that a delayed exhausted response cannot recreate a block after newer successful use. Inherited success preserves selection history.
- Restart reconstructs quota state from existing storage. Configuration and policy reload retain existing snapshot ownership; no separate quota cache or new persisted state is introduced.

- Saved-only clearing is observed separately from natural expiry, successful inference, and login/logout. A separate process removes only the selected profile's five persisted block fields, preserving credentials, configuration, and all other usage fields. On the unchanged base, explicit `models.authStatus({agentId:'main',refresh:true})` followed by the next ordinary request succeeded on the same running Gateway while the prior health observation was still recent and no five-minute advance occurred. Actual command-line status returned exit 0 with no unusable entry for the selected profile. Both base browser cases also displayed Ready on the provider header and profile; their retained screenshots were inspected.
- The refresh path is `src/gateway/server-methods/models-auth-status.ts` → `src/secrets/runtime.ts` → prepared/live bookkeeping activation in `src/secrets/runtime-state.ts`. Browser reads then pass through `ui/src/lib/model-auth.ts` and the provider-page load/derive/render owners above. The correction run passes next-turn, command-line, and browser observations against its exact source manifest. It also requires the actual inference bearer credential and account; the earlier base pass did not contain that strengthened assertion. A refresh acknowledgement or responsive `models.list` is not a browser result.

## Tests

The four observed response races are covered through registered Gateway requests, with the provider response held until the competing operation settles:

| Race | Regression and registered entry points | Required last-consumer result |
| --- | --- | --- |
| Older exhausted usage after newer main success | Committed `quota-reset.e2e.test.ts` stale-success case; `sessions.companion.ask`, `chat.send`, `agent.wait`, `chat.history`; independent overlap also recorded | Immediate follow-up chat succeeds; old exhaustion cannot recreate quota |
| Older exhausted usage after newer inherited-child success | Independent recorded regression; real `sessions_spawn`, then child `chat.send` and `agent.wait` | Immediate child follow-up succeeds on the same inherited account; selection history stays unchanged |
| Older exhausted usage after newer available usage | Independent recorded regression; `sessions.companion.ask` through failure accounting, followed by `chat.send` | The older response cannot recreate quota after the newer available observation; primary chat succeeds |
| Older available usage after newer exhaustion | Independent recorded regression; `sessions.companion.ask` overlaps native quota failure through `chat.send` | The newer persisted quota remains; immediate primary follow-up stays blocked |

The inherited and opposite-observation overlaps are retained end-to-end acceptance regressions, not additional committed test cases. Credential replacement during a held response is a separate recorded control.

- New registered regression: `test/e2e/qa-lab/runtime/quota-reset.e2e.test.ts` starts the full Gateway and native app-server, drives `chat.send`, waits through `agent.wait`, and reads the result through `chat.history`. Provider responses create the block; no quota rows are injected. Both quota sources cover restored capacity, still-exhausted capacity, early retry, and revoked credentials. The native-source case warms three sessions, then receives concurrent real quota failures while usage returns HTTP 503, exposing competing local backoff. Additional controls cover extra buckets, workspace restrictions, spending limits, malformed usage, and expiry through existing renewal.
- The fixture preload redirects only named upstream endpoints and advances Gateway wall-clock reads. Real timers and the native process clock remain unchanged. A separate baseline run waited five real minutes and drove registered chat, `cron.add`/`cron.run`/`cron.runs`, and a child actually created through `sessions_spawn`.
- Removed unit test `leaves non-WHAM blocks outside the half-open probe path`: retired source exclusion; replaced by the registered native-source case at `test/e2e/qa-lab/runtime/quota-reset.e2e.test.ts:68`.
- Removed unit test `does not re-probe a WHAM block inside the half-open interval` and its mirrored 45-minute constant: retired interval; replaced by registered early-retry and five-minute recovery cases at `test/e2e/qa-lab/runtime/quota-reset.e2e.test.ts:68`. Existing newer-failure, auth-classification, and tie-priority tests remain.
- Existing usage persistence fakes expose their committed store to the owner read; no new production test seam is added. The three exact rewritten unit cases below consume the returned Promise while preserving concurrent starts, held-response ordering, and all assertions. Their direct helper calls are owner-level tests, not registered Gateway entry points.

| Exact rewritten unit test | Owner-level purpose | Actual registered evidence counterpart and limit |
| --- | --- | --- |
| `src/agents/auth-profiles/usage.test.ts:1322` — “half-opens a stale long WHAM block and clears it when capacity returns” | Starts two requests before awaiting both; one provider read clears the block, records the observation time, and performs claim/result transactions. | Committed `quota-reset.e2e.test.ts:68` recovery rows drive `chat.send` → `agent.wait` → `chat.history`. Retained concurrent main/child/registered-cron recovery additionally proves shared pending work: one held usage request serves all three terminal consumers. The unit's two-lock count remains a local owner assertion. |
| `src/agents/auth-profiles/usage.test.ts:1354` — “re-arms a stale WHAM block from the latest blocked snapshot” | Awaits an exhausted usage observation; persists the new reset and keeps its model scope and observation time. | Committed `quota-reset.e2e.test.ts:310` exhausted-demand branch uses the same registered chat methods and requires denied output with persisted future blocking; additional denial controls at `:277` forbid inference. These are final-consumer counterparts for retained exhaustion, not proof of the unit's exact one-hour fixture timestamp. |
| `src/agents/auth-profiles/usage.test.ts:1392` — “does not apply an available result over a newer WHAM block” | Holds an older available response, writes a newer failure generation, releases the old result, and checks that the newer block survives. | Retained registered race starts older `sessions.companion.ask`, then creates newer native quota through `chat.send` and waits with `agent.wait` before releasing the older response. The persisted newer block remains and the immediate primary follow-up is refused. This exercises the common generation invariant through failure-accounting/ordinary-chat callers; it is not mislabeled as a committed standalone admission-race test. |
- The baseline typed-reference census also found unchanged fixtures for expiry, round-robin ordering, inherited persistence, runtime snapshots, state publication, profile replacement, session preferences, auth rotation, and final auth-plan selection. These references identify sibling contracts; they are not a claim that all those suites were rerun.
- The final census expands the historical 80-path inventory to 112 paths, with 648 resolved queries out of 654 and 5,783 references. All original paths remain. The three historical SDK-agent-runtime gaps and three new SDK-simple-completion gaps remain explicitly named in Consumers; no gap is treated as zero consumers. Dynamic personal-account persistence, normalization keys, aliases, native callers, and additional approval/worker/speech boundaries retain their source dispositions.
- The additional registered case exercises `sessions.companion.ask` with an ordinary utility-model 429 and a fresh available usage response, followed immediately by `chat.send` on the primary model. It checks primary recovery and the utility cooldown at the inference boundary before successful-use cleanup.
- The sixth registered case starts without a quota block, holds an older exhausted usage response, completes a newer main chat, and requires the immediate follow-up chat to succeed after the old response arrives. The success-owner timestamp prevents the delayed response from recreating the block. Independent proof also covered a genuinely spawned child on the same inherited account.
- An older-host upgrade matrix is not applicable: no plugin, SDK, updater, Doctor, schema, or version surface changes. Existing captured compound quota/cooldown state recovered without startup rewrites. Restart proof before the final rebase preserved the quota row through a whole-state copy and startup, then recovered on ordinary demand after restoring the test clock; this is virtual-time supervisor-completion proof.

The storage-failure regressions also use the registered `chat.send` → `agent.wait` → `chat.history` flow, after real provider quota responses establish the block. Separate claim-write and available-result-write cases use an operating-system file-size limit and an oversized SQLite write to produce native IOERR_WRITE (`778`). A connection-local TEMP trigger targets the actual auth-state transition; the persisted schema, normal schema checks, statement cache, and transactions remain unchanged. Direct backup calls before and after the fault must succeed. The affected ordinary turn must use that backup, make no primary inference request, and retain the original quota. A third case raises a native constraint error (`1811`) and requires the ordinary turn to fail. Each case disarms the fault and then requires ordinary primary recovery; the in-flight entry cannot remain stuck. Windows skips these POSIX-only cases.

Shared test provider, clock, Gateway startup, chat, and persisted-state observation code moved to `quota-reset.test-support.ts`; all six existing scenarios and assertions remain. No production test hook was added. The preload records native errors and rethrows the same objects. Initial fixture attempts exposed and retained two setup failures: a file-size signal killed the CLI child, and a persistent trigger correctly failed canonical schema validation. The final fixture handles the signal across CLI respawns and uses only a TEMP trigger.


The correction tests add these explicit final-consumer cases:

| Case | Registered or actual consumer | Current evidence state |
| --- | --- | --- |
| Host isolated recovery and exhausted refusal | `test/e2e/qa-lab/runtime/quota-reset.isolated.e2e.test.ts` invokes `tools.invoke` for the configured isolated task plugin and checks returned structured text | The unchanged base fails restored isolated recovery. The tested common-owner correction passes delivered isolated recovery, early refusal, and due-exhausted refusal without an intervening ordinary recovery turn. |
| Credential deletion during an isolated probe | The same registered task path holds an observation, deletes the exact selected profile through its owner, and checks pinned denial plus explicit alternate success | Passes on the recorded correction sources: the pending and fresh pinned task are refused after removal, while an explicitly selected healthy alternate succeeds. This is the actual isolated route, not companion/child proof. |
| Automatic recovery status | `ui/src/e2e/quota-reset-status.real-gateway.e2e.test.ts`: ordinary chat, actual `models status --json`, and the real built provider page | Passes on the recorded correction sources: exact inference bearer/account, successful chat, command exit 0, empty `unusableProfiles`, OAuth status `ok`, and provider-header/profile `Ready`. Final screenshots were inspected. |
| Recent saved-only clear and refresh | Separate-process manual clear → `models.authStatus({refresh:true})` → next ordinary chat before later command/browser reads | Passes on the recorded correction sources without restart or five-minute advance: exact inference bearer/account, successful next turn, command exit 0, empty `unusableProfiles`, OAuth status `ok`, and provider-header/profile `Ready`. Final screenshots were inspected. |

The external clear is owned by `test/e2e/qa-lab/runtime/quota-reset.saved-clear.test-support.mjs`. The browser uses the actual Gateway UI, canonical invitation dismissal, and the existing retained-artifact allocator. No mocked RPC or substitute renderer supplies the status. The command-line read is a separate process. These are observation tests; no new quota/state writer is introduced into production.


The existing simple/isolated test adjustments preserve their prior contracts. Five call assertions add only the trusted lifecycle callback as the second argument; no original expected field, dispatch denial, result, or lease assertion is removed.

| Existing rewritten test | Retained purpose and registered-proof relationship |
| --- | --- |
| `src/agents/isolated-completion.test.ts` — `captures call-owned choices and authority before admission (retired: %s)` | Captured model/profile/stream choices survive caller mutation; a retired owner cannot prepare or dispatch, and the lease is released. Focused ownership evidence; the registered C1 task supplies the actual host-preparation path. |
| `src/agents/isolated-completion.test.ts` — `uses admitted config and directories for a newly owned %s completion` | Host and command execution use admitted config/directories and release the lease. C1 traverses host preparation; the command branch retains separate unit coverage. |
| `src/agents/isolated-completion.test.ts` — `passes one prepared route to the selected harness and returns text` | Legacy host result, selected prompt/fingerprint, and prepared-owner lifetime remain checked. C1 host V2 recovery is not presented as legacy-harness execution. |
| `src/agents/harness/isolated-completion.native-auth.test.ts` — `rejects a retired native route before dispatch (API sibling: %s)` | Retirement blocks native dispatch; only an already prepared runnable API sibling can receive host authorization. This is retained route-authority evidence, not registered quota proof. |
| `src/agents/harness/isolated-completion.native-auth.test.ts` — `uses host authorization for V2 API-key routes` | API-key host authorization and resource lifetime remain checked. C1's quota-bearing host V2 credential does not replace this separate API-key contract. |

`src/agents/simple-completion-runtime.test.ts` adds a no-op mock for `reconcileAuthProfileQuotaBlocks` to its existing model/auth wiring suite. Its intentionally partial store mock is not a live database owner. Existing token exchange, mixed-credential, exact-profile, route materialization, missing-auth, and provider-hook checks remain; the suite is not described as quota recovery proof. Real quota behavior belongs to the registered C1 and retained ordinary/storage cases. These adjustments add nine test lines and remove no test or assertion.

The C1 registered task specifically exercises host V2 preparation. No legacy-harness, API-key, direct-provider, local-inference, speech, or every-plugin output runtime pass is inferred from sharing that owner.

The CI follow-up preserves the existing guards and adds the new case to their expected inventories:

| Existing test / entry point | Preserved behavior |
| --- | --- |
| `test/scripts/ci-workflow-guards.test.ts` — `keeps private Control UI servers and resource-sensitive files under one serial owner` | The source parser now recognizes the canonical quota fixture's own Gateway; discovered and declared private-server inventories still match exactly. |
| Same file — `selects the complete real-Gateway command without retrying failures (frozen: $frozen, prebuilt: $prebuilt, exit: $childExit)` | Its shell-adapter checks now see the test in the actual workflow command; missing current config and child exit 42 still fail without retry. |
| `test/vitest-ui-e2e-config.test.ts` — `owns the complete inventory once and shards the project union without losing QA Lab or real-Gateway siblings` | Real Vitest discovery and shard/project unions include the new file without duplication or lost siblings. |
| Same file — `admits every prebuilt real-Gateway file once with the native worker cap %s` | Both cap cases require the new file in serial standalone phase 1, with one worker and no file parallelism. The audited parallel phase is unchanged. |

The two browser rows retain all prior assertions. The saved-clear receipt now explicitly requires the selected profile's usage before enumerating it, so a missing row fails rather than becoming a dummy value. This resolves the strict indexed-access type error without a cast or fallback.

## Evidence

- A local recording HTTP/WebSocket endpoint supplied quota and reset responses to the full Gateway and native client. Their normal failure, persistence, and admission paths produced the observed state.
- Pinned-base full Gateway reproduction: both quota sources reject restored capacity after five real minutes on chat, registered cron, and a genuinely spawned child. Still-exhausted and revoked-account controls were recorded separately.
- Before the common-preparation correction, the storage-fallback source passed 261 owner/sibling tests across six files and all nine full Gateway cases after a fresh runtime build. Its independent source review found no actionable issue. These retained results do not claim final C1/C2 runtime coverage.
- The earlier production revision passed 332 tests across nine owner/sibling files and six full Gateway cases, including both quota sources, fresh-usage reconciliation, cross-model preservation, delayed-response races, token renewal through its existing owner, and agent-selection/rotation siblings. The later test-only Promise correction passed 112 existing owner/inheritance tests; production code is unchanged.
- Independent proof before the final rebase passed ordinary main, genuine-child, and registered-cron quota recovery with a virtual five-minute advance. One fresh usage request served all three. Separate earlier baseline/candidate runs used real five-minute waits; the time models are recorded separately.
- That independent proof also covered both stale-observation directions, main/child success races without prior quota, actual `/btw` with automatic backup selection, config hot reload, preserved-state restart, and credential replacement while a response was held. A revoked native account failed authentication without creating a quota block. The quota implementation stayed unchanged through the rebase; `/btw` also retained two upstream agent-selection arguments. The focused and full-Gateway tests above ran again on the retained pre-C1 integration.
- Storage fallback regression: both claim and result failures reproduced on the earlier PR head through ordinary chat, with native IOERR_WRITE and successful independent backup controls; the failure was at the final ordinary-chat result assertion. The corrected source passed both I/O cases and the fatal-constraint control; all three also recovered after disarming.
- Max-lines and assertion safety preflight passed against the pinned base. CI merge-tree structural gates remain required.
- Telegram Test Server proof on revision 7 completed with runner exit 0. Ordinary chat and `/btw` each recovered from a separate provider-created block after a measured real five-minute wait, with delivered replies, matching native receipts, and cleared quota state. Preserved live captures and the terminal runner result are accepted proof. The final archive was lost when the runner host ended; retained observations and cleanup reconciliation are recorded separately, without claiming the missing archive exists.
- A subsequent read-only Telegram observation retrieved all six original replies with no recorded edit, verified the preserved capture integrity, and completed its own credential cleanup. It sent no messages and made no inference request. Historical process and original credential-release receipts remain unavailable.


- The combined runtime and Control UI build completed successfully. The retained correction run then passed 11/11 Gateway cases (nine existing cases and two registered isolated cases) and 2/2 real-browser cases. Both test commands returned exit 0; the native run completed with exit 0 in 3m36.056s.
- Focused simple/isolated coverage totals 93 passing tests across six files: 26 retained passes in selection (13), generation (8), and isolated resources (5), plus 67 passes in the corrected isolated-completion (33), simple-completion-runtime (22), and native-auth (12) rerun. The first run's argument-shape and incomplete-mock failures remain recorded; the callback expectations and model/auth wiring mock changed before the rerun. Counts do not add overlapping passes twice.
- C2 base v4 passed both real-Gateway browser rows: automatic recovery and recent saved-only clearing. Each observed a successful next ordinary turn on the same running Gateway, command-line exit 0 with no unusable profile, and actual provider-header/profile Ready status. Screenshots were inspected. This closes the original base evidence omission. The later correction run additionally passes the strengthened inference-account assertion; that guarantee is not retroactively attributed to the earlier base run.
- Review found and corrected a C2 assertion gap: a success marker and Ready status alone did not prove the inference account because the fixture accepts arbitrary tokens. The final test checks the exact bearer credential and account before later status reads, and both correction browser cases pass that check. The earlier base pass remains evidence for its original assertions only.

- Retained product-correction proof identity: signed head `b6ec779939db15a7ffd56aba4bb539d1ea9a2f7d`, tree `7c8e35e475cdf4ed7b4fc92a36b4bc3c6d8a36db`, matches every file in the successful 11-case Gateway and two-case browser run's source manifest. That signature was verified. The final test/CI-only follow-up is bound below. The retained unchanged `f1409c01426dae3b26ab4b855115602ccfe97784` isolated failure remains the regression baseline; the correction's pass is not attributed to that old head. The source-manifest digest is `8c287e8643029ece83d64843354758e37b5a136a3c9c8649eccfd15a8667e45d`; the complete run archive digest is `4e5477bcd4bf30f1169005508fcf67927c93ff1532becfc7407a690a8afd4c04`. Full logs, build identity, per-file source hashes, browser readbacks, and captures are retained.
- Final census evidence is also retained: candidate query output digest `d6b7fd0e9bb275a7aa070bb619e38f328d42116f82903a4b0c3f68fa7f083dbb` and complete base/candidate summary digest `4c6703cac89a1d353b36436f6a970baf16a48ebc433709c49a5a150af231dee7`. Its 32-path delta is fully dispositioned without relabeling approval decisions, placed-worker execution, or speech/summary output as isolated-task runtime proof.

- Test/CI follow-up identity: signed head `7739004cef2968e27515926c8afd0d85f7800ada`, tree `77576e488108c624a8eb977c53772475dcfc1c45`. Its five-file diff contains only the required-row check and complete test registration. All product source matches `b6ec779939db15a7ffd56aba4bb539d1ea9a2f7d`. On the matching product build, 43 partition tests, six selected workflow cases, and both browser cases passed; formatting, focused lint, and both ratchets passed. The native phase exited 0 in 1m26.627s. Complete proof archive SHA-256: `bc8d884d06e083c410db698ac6218a85724411f1dfff56df5ab4d61ac11532b2`.
- The preceding actual CI merge `38e5d3cf5f68859d2c42e8fa563a3ecef3b1810c` passed all four structural commands, the 654-pair typed census with the same 112 paths and six gaps, a matching combined build, both isolated cases, and both browser cases. Its three CI failures were the indexed-access error and the two test-inventory omissions addressed by this follow-up. Fresh final-head CI and merge-tree checks remain required; no unrelated-red exception is claimed.

- Final browser-launch follow-up: `625c2c137262c475e1bec74e2ede9287672116f0` only consumes the browser executable supplied by the existing UI test setup. Both complete browser cases pass with an empty default browser cache and an explicit provisioned executable, preserving every quota/account/state/status assertion. Native exit 0 in 34.750s; proof archive SHA-256 `80b0ef1e5b2f069a758f59a2b203beb0ad56f61cdda344bd0c3580fcce4eb31b`. The preceding CI failure reached browser launch and requested a missing default headless binary; it was test setup, not a quota product failure. Product source remains byte-identical to the earlier correction.
- Retained native-source provenance is now bound: all three inspected quota/error/window source files match clean committed blobs at `e5769939113536eb72752660bf7d1903f799d198`, including their original retained hashes. This identifies the source snapshot; it does not assert that commit built the separately version-pinned native package exercised by the Gateway tests.

## Required publication checks

- Publish the signed, tested head and verify that the remote head matches it.
- After publication, complete the required exact-head review and full CI checks, including the actual integration identity and applicable maintainer/merge authority. The successful local build, 93 focused owner tests, 11 Gateway cases, and two browser cases do not waive those gates or claim they are already green.

The missing isolated host branch, saved-clear/status observations, and exact consumer/unit mappings were visible in earlier reviews. They are recorded as reviewer misses, not additional worker convergence strikes. This body does not grant merge authority.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-12 18:20:19 +05:30
Peter Steinberger
f9ccd07ded
fix: skip official setup approvals and default to Astra (#145646)
* fix: skip official setup approvals and default to Astra

* test: align default-model expectations across consumers

* test: align attachment catalog with Astra default

* fix: preserve configured Codex catalog model selections
2026-09-12 00:34:25 -07:00
Dallin Romney
1ac85670b6
fix(ci): recover generated PR push timeouts (#144878) 2026-09-12 00:24:12 -07:00
Peter Steinberger
1754576c2c
ci: use measured 16-class runners for current artifact builds (#145611) 2026-09-11 23:26:34 -07:00
Peter Steinberger
b5134de70a
feat: manage and reload plugins without Gateway restarts (#145484)
* feat: manage and reload plugins without Gateway restarts

* fix(plugins): preserve Bun loading and align lifecycle fixtures
2026-09-11 23:12:58 -07:00
Peter Steinberger
d5e6520990
fix(ci): report ClawSweeper dispatch API failures truthfully (#145617)
* fix(ci): report ClawSweeper dispatch API failures truthfully

* test(ci): include ClawSweeper behavior suite in workflow routing
2026-09-11 23:04:34 -07:00
Peter Steinberger
01e00e442f
feat(linux): add Omarchy agents panel with desktop handoff (#145593)
* feat(linux): add Omarchy agents panel with desktop handoff

* fix(omarchy): preserve Unicode session selection after prompts
2026-09-11 22:09:06 -07:00
Peter Steinberger
ce5c1cb1b1
feat: support inline browser panels in the Tauri companion (#145572)
* feat: support inline browser panels in the Tauri companion

* test: cover inline browser proof in companion CI routing

* fix: register the native browser validation entry point
2026-09-12 04:29:44 +00:00
Hannes Rudolph
2227743f74
refactor: split release changelogs and synchronize docs mirrors (#145464)
* refactor: split release changelogs and synchronize docs mirrors

* fix: complete split changelog instructions and validation wiring

* fix: complete release changelog mirror integration

Regenerate existing docs mirrors within the docs-agent publication boundary, preserve one HTML release heading, and package links for oversized mirrors without changing frozen records. Update release publisher and test-routing fixtures for the shared changelog resolver.

* test: align docs agent Git ownership fixtures

Keep failure injection aligned with staged-index validation and mirror staging. Preserve native Git producer exit codes and verify both cached-index producers without weakening process-drain assertions.

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-11 21:19:18 -07:00
Vincent Koc
dda6470454
feat(desktop): match virtual worker displays to the viewport (#143938)
* feat(desktop): match virtual worker displays to the viewport

Keep Fit as the default and gate Match on provider-owned resize permission and authenticated controller ownership. Preserve sizing choices during reconnects and retire resize authority with input control.

Pass resolved SSH VNC credentials through optional auth metadata and reconcile lost physical modifier releases after native popups. Add instrumented production Gateway/XFCE resize proof and native selection regressions.

* fix(desktop): align renderer exports and suppression inventory

* test(desktop): cover real node carrier resize qualification

* test(workers): align repository access provider fixture

* fix(ci): run real desktop resize proof with an upstream fixture

* fix(ci): bind desktop proof to raw merge identity

* fix(ci): retain safe desktop proof failure diagnostics

* fix(ci): expose safe desktop SSH setup diagnostics

* fix(ci): keep desktop proof phase tuple private

* fix(ci): retain desktop proof phase after timeout

* fix(ci): observe Gateway startup at desktop proof timeout

* fix(ci): run desktop proof against built Gateway

* fix(ci): assert visible desktop recovery state

* fix(ci): remove retired desktop spy allowance

* test(ui): model targeted desktop status in sizing fixtures
2026-09-12 11:26:46 +08:00
Peter Steinberger
f89fa78584
feat(radius): connect models with browser sign-in and native streaming (#145377)
* refactor(channels): simplify progress rendering ownership

Keep visibility in the compositor and remove redundant mode flags from internal layout helpers. Separate Slack attention-title formatting from native approval task identity construction, avoiding discarded hashes and task objects in Block Kit rendering.

Preserve public SDK and output contracts. Remove one duplicate test left after native start/update builder unification. Related: #145345, #140037.

* fix(ui): let the popover own initial picker focus

WebAwesome already focuses the autofocus input before the opening animation. Remove the duplicate after-show focus that stole focus from a newly opened cloud configuration card. Replace the handler-only unit expectation with a deterministic browser regression that controls the actual animation and preserves initial focus, cloud focus, and machine selection.

* feat(radius): add native model provider and browser sign-in

Support Radius organization API keys and device OAuth, discover account-visible models and tiered prices, and stream native Pi messages through the provider plugin contract. Add setup docs and meaningful auth, catalog, and tool-stream coverage. Regenerate the standard plugin config inventory for the new provider.

* refactor(radius): satisfy provider validation guards

Require captured OAuth and model requests in strict test types. Project optional stream fields directly while preserving validation of supplied values. Focused tests and independent review pass.

* fix(radius): accept live fractional OAuth polling intervals

Radius device authorization advertises interval=0.1 seconds. Accept finite positive fractional intervals and round milliseconds upward; retain expiry and timer bounds. The regression failed against the original parser and all 20 OAuth tests now pass. Live pairing reaches browser authorization.

* fix(radius): publish native model routes and resolve cold starts

Declare pi-messages in the shared API data contract and emit it from authenticated Radius catalogs before registry validation. Resolve cold-start models through the existing provider hook using the same account catalog and pinned auth profile. Remove late normalization hooks and cover real registry deserialization, schema admission, and account-scoped resolution. Live authorization and catalog discovery pass; inference reaches Radius and reports the account billing gate.

* build(radius): integrate plugin publication and documentation metadata

Document the plugin package, exclude its external dist tree from the core npm package, and update extension label routing and the exact release inventory. Regenerate the plugin inventory/reference pages and brand glossary entries. Publication/build-selection and labeler coverage tests pass.
2026-09-11 19:29:25 -07:00
Vincent Koc
f1e430a888
fix(ci): resolve performance targets with split config schemas (#145366) 2026-09-12 07:18:17 +08:00
Peter Steinberger
6cc8bb1a4b
perf(build): overlap staged SDK compilers within memory capacity (#145090)
* perf(build): overlap staged SDK compilers within memory capacity

* test(ci): align hybrid real-Gateway runner expectation

* test(build): keep nested writer fixtures on real plans

* fix(build): retain both joined compiler failures
2026-09-11 13:11:01 -07:00
Peter Steinberger
0d3f4501fd
ci(android): retain existing test timing reports (#145131) 2026-09-11 12:36:47 -07:00
Dallin Romney
8bdd9807e6
fix(e2e): qualify rich Telegram scenario by source (#144841) 2026-09-11 12:25:25 -07:00
Peter Steinberger
7a92d3633b
ci: credit the required max-lines ratchet in test routing (#144359) 2026-09-11 10:15:41 -07:00
Peter Steinberger
5ff271b457
ci: avoid duplicate startup corpus runs on main (#145078) 2026-09-11 10:14:15 -07:00
Peter Steinberger
6cf57708a0
chore(ci): remove the retired credential scanner (#145053)
* chore(ci): remove the retired credential scanner

* docs: remove retired scanner reference from release notes
2026-09-11 09:26:43 -07:00
Peter Steinberger
bbeedd9a0d
fix(release): enforce trailing-completed-month extended-stable rule (#143679)
`docs/reference/RELEASING.md` has defined extended-stable as "the trailing
completed month's `.33+` maintenance line" since #99352, but the guard only
required protected `main` to be in *any* later calendar month. A 2026.6.35
publish therefore reached the publish step with `main` on 2026.9.3 — three
months stale, on a line policy says was already retired, against a public
promise of one monthly line.

Require the month gap to be exactly one, so an older `.33+` line retires when
`main` advances another month. `BYPASS_EXTENDED_STABLE_GUARD` already
short-circuits above this check and stays the sanctioned exception for an
explicitly approved emergency or retirement release; no new input is added.

The rejection message names the version that would be legal and the bypass, and
the docs, workflow input description, and root AGENTS.md now state the rule the
code enforces.
2026-09-11 08:59:47 -07:00
Vincent Koc
5ff5dcefeb
fix(perf): isolate external evaluator from candidate runtime (#144162)
* fix(perf): isolate benchmark subprocess transport

Run transported CLI samples and their temporary state through the explicit SUT transport while preserving native measurements. Reject unsupported transported runtime RSS and profiles before launching candidate code.

Punchcard-Session: coral-orchard-orchard-d4

* fix(perf): isolate external evaluator from candidate runtime

Add the standalone off-Actions harness with separate runner and candidate identities, bounded collection, and preserved workload receipts. Retain custom evaluator diagnostics without granting gate authority; leave workflow activation to the dependent change.

Punchcard-Session: coral-orchard-orchard-d4

* fix(perf): reject mixed benchmark execution modes

Record evaluator-owned execution identity per suite and enforce compatible modes in comparisons, budget checks, and source summaries while preserving independent RSS semantics.

Punchcard-Session: coral-orchard-orchard-d4

* fix(perf): retain receipts after custom diagnostic failures

Capture the complete custom collection, validation, and summary phase without disabling errexit. Preserve workload failure precedence and always reach the separate sealed-export and receipt finalizer.

Punchcard-Session: coral-orchard-orchard-d4

* test(ci): preserve fixed IMDS probe argument tuples

Punchcard-Session: coral-orchard-orchard-d4
2026-09-11 21:09:08 +08:00
Ayaan Zaidi
0e9391f65d
fix(ui): retain controls after partial catalog refresh (#144750)
Fixes #144726.

## What Problem This Solves

Fixes an issue where users lose effort and speed controls in existing chats when another provider fails to refresh its model catalog. The operator reported this on 2026.9.4; New Session retained its controls under the same condition.

## Why This Change Was Made

Both pages now share catalog readiness and retain partial-refresh warnings separately from failed requests. Controls use the selected model's capabilities and availability, while the model menu keeps the refresh notice visible.

## User Impact

| Same usable catalog with a partial refresh warning | Before | After |
| --- | --- | --- |
| Existing chat | Models listed; effort and speed hidden | Models and supported controls remain available |
| New Session | Supported controls available | Supported controls remain available |
| Warning | Catalog error presentation | Notice without changing readiness |

## Evidence

- Pinned-main mock and real-Gateway browser reproductions fail on the missing existing-chat effort control; candidate passes both pages.
- Eight browser cases cover partial provider failure, warning recovery with retained settings, locked selection, empty catalogs, rejected reads with and without retained rows, unavailable selected models, and non-reasoning models.
- Config/state corpus: 50 passed. Extended owner/default tests: 127 passed. Browser regressions and affected siblings: 11 passed. CI inventory checks: 40 passed.
- Independent browser acceptance passed. Telegram Test Server acknowledged the native effort command.
- Focused formatting/lint and diff checks passed. No schema or protocol change.

The real-Gateway fixture uses fake credentials and a simulated Copilot HTTP 503, with no network inference. The captures below are cropped to the complete conversation area; the unrelated sidebar is excluded from both baseline and candidate images.

| Surface | Before | After |
| --- | --- | --- |
| Existing chat | ![Existing chat before](https://github.com/user-attachments/assets/403a3fb3-0edf-47f1-98f3-28793794ee3e) | ![Existing chat after](https://github.com/user-attachments/assets/7389d8f4-86e9-44fa-a63c-a9d271a2fc6c) |
| New Session | ![New Session before](https://github.com/user-attachments/assets/2730b548-02a4-4486-a5be-e548d00ee075) | ![New Session after](https://github.com/user-attachments/assets/9281e132-d8c5-4ddf-af0c-474696346df7) |

Refresh notice remains visible beside the supported controls:

![Refresh notice with working effort control](https://github.com/user-attachments/assets/8899cfe4-b296-42d3-9a25-2e9a254414ac)

Known unrelated main failure: the session-link presentation test expects `inline-grid` but observes `grid` in an existing flex container. The same test and styles fail in [main CI](https://github.com/openclaw/openclaw/actions/runs/34567817345/job/103163547116) and are unchanged by this PR.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-11 13:10:52 +05:30
RoboClaw
088d0f5b87
fix(release): publish frozen plugins with valid ClawHub categories (#144615)
* fix(release): bind ClawHub categories to trusted tooling

Project reviewed single-category metadata through the existing isolated package
manifest overlay while retaining the frozen candidate version, runtime files,
package manifest, and generated channel configuration. Pack with trusted release
tooling before the unchanged transaction-sealing and approval boundaries.

Use existing reviewed category assignments; retain the shared runtime vocabulary
for acpx and codex. Do not change published npm artifacts or the final release tag.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

* fix(release): keep category validation native-node safe

Extract the unchanged category contract into a dependency-free module so
pre-build packaging and updater entrypoints do not load compiled imports.
Preserve existing barrel exports and validation behavior.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

* test(transcripts): keep capture lifecycle fixture model-free

Explicitly disable utility routing in the capture replacement fixture so
its real heuristic-summary and lifecycle assertions do not depend on a
provider metadata snapshot retained by an earlier test file. Keep all
assertions, deadlines, and production code unchanged.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-11 00:22:21 -07:00
Ayaan Zaidi
2128a60815
chore: gate startup upgrades against prior-release state (#144659)
Related: #144626.

## What Problem This Solves

The config startup corpus uses empty state. It cannot detect an upgrade that loses existing sessions, auth profiles, cron jobs, or Control UI preferences, or leaves a prior-release SQLite schema behind.

## Why This Change Was Made

Add generated state snapshots from stable 2026.9.2 and frozen 2026.9.3 source, then cross them with the config corpus through the existing Doctor and Gateway startup owners. Check preserved session text, credential payloads, cron behavior, preferences, current schema versions, and database integrity after two repair passes. Add the five distinct sanitized operator config shapes instead of relying on two consolidated examples. Keep the combined container fixture’s built-in references resolvable so its wildcard allowlist still exposes models.

The Linux-generated snapshots retain POSIX cron partition keys. The state matrix skips Windows; fixture inventory still runs on every platform. Windows upgrade coverage requires Windows-generated state. The fixtures contain only generated data and fake credentials. Production code and schemas are unchanged. Both selected releases already use SQLite sessions and shared auth storage; this corpus tests their upgrade and preservation contracts, not older JSON transcript imports or external CLI authentication.

## User Impact

No runtime behavior changes. Future changes gain a required upgrade-state check beside the existing config corpus check.

## Evidence

- Both required release snapshots passed all 15 config combinations, preserving session identity/text, fake profile payloads, complete cron behavior, preferences, target schema versions and SQLite integrity after repeat Doctor repair.
- The complete 56-test corpus/sibling selection passed on an isolated Testbox. After the CI correction, all 16 config tests and the exact unused-file scan passed again. The three affected container-fixture cases and its real Gateway flow also passed after correcting inherited placeholder references. A deliberate wrong expected session ID failed the preservation assertion; the original manifest was restored.
- Real Gateway proof passed for all 15 config shapes with frozen state, plus previous-stable state with `operator-host`: readiness, `models list`, retained history, completed `/status` turn, and history after resume. Outbound channels, cron and browser services were disabled; no inference was requested.
- Strict read-only Doctor completed with exit 1 for fixture warnings (loopback binding and fake plaintext credentials). CrossClaw and Peanutto also report the intentionally unreachable `.invalid` tool-server URLs. These diagnostics remain visible; no production suppression was added.
- The full 31-test Linux state suite passed again after the platform-scope correction. Post-shutdown SQLite integrity also passed across the retained Gateway scenarios. Formatting, focused lint, workflow checks, assertion safety, and file-size guards passed. Runtime proof used intake main `9dca5870c6`; source review confirmed the final rebase leaves the exercised paths unchanged.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-11 10:56:46 +05:30
Peter Steinberger
b7402c5b9e
ci: preserve tooling metadata in precise test plans (#144621) 2026-09-10 21:24:17 -07:00
Ayaan Zaidi
938a072ebb
test(config): retain operator startup configuration corpus (#144626)
## What Problem This Solves

A provider entry without custom model rows passed synthetic coverage and then crashed a Gateway. This maintainer-requested corpus makes real operator configuration shapes part of pull-request validation.

Related: #144599

## Why This Change Was Made

The suite loads two sanitized operator configurations and eight canonical shapes through shared Doctor normalization, the Gateway startup snapshot owner, prepared static catalog construction, and provider login-choice resolution. It preserves absent optional model lists so normalization cannot hide the startup regression. The existing core CI checks run this corpus on core changes.

## User Impact

No production behavior changes. New startup regressions in retained operator config shapes fail CI before landing.

## Evidence

- The missing-list cases fail on the baseline at the same registry call as a real Gateway startup.
- Final rebased head: 13 focused tests pass, including the corpus and a legacy-provider sibling. The same command used by CI passes.
- All ten fixtures pass real Gateway startup, CLI listing, full-inventory RPC, webchat command responses, and browser rendering. The multi-agent fixture also exercises both agents.
- Telegram Test Server proof covers the provider picker, owner login choices, and non-owner rejection with zero model requests.
- Corpus runs use synthetic credentials and loopback-only network namespaces. Private plugin configuration uses an inert fixture plugin.
- Legacy API aliases require `doctor --fix` before direct startup; recovery is covered.
- The requested local source config was absent. Both available remote operator sources are retained, with structure-preserving sanitization.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-11 09:26:24 +05:30
Ayaan Zaidi
09e2b15466
fix(models): share API-key editing and removal (#144420)
Related: #136257
Supersedes #142851. Builds on #144181.

## What Problem This Solves

Fixes an issue where API keys saved on the Models page followed a different persistence path from CLI keys, and CLI logout refused to remove profiles referenced by provider configuration.

## Why This Change Was Made

Models and CLI key operations now share credential selection, persistence, and removal. Removal clears references for the same captured credential plan that the store deletes; concurrent replacements are preserved. Both paths use the Gateway auth-refresh owner and retain recovery guidance after a committed write.

If another client changes the provider binding while the CLI key prompt is open, the saved key does not replace the newer binding. The CLI reports the committed key and tells the user to reopen Models. Internal write controls stay outside the public plugin interface.

Reference-backed profiles cannot be converted to inline keys by replacement. If a concurrent credential generation rejects removal after config cleanup, the remover restores only cleanup-owned config values that no later writer changed.

If removal throws or commits only some owner stores, the removal owner reads the targeted stores again. It restores references only for credentials that still exist and keeps successful deletions removed.

Recovery replays only the captured config delta for targeted survivors. A retained token or external-secret profile keeps its provider binding.

## User Impact

- Edit or remove saved API keys through Models or the CLI with the same stored result.
- Keep model defaults, connection settings, account metadata, credential copy restrictions, tokens, and reference-backed keys intact.
- Explicit CLI profile selection still supports named backups without changing the active connection.
- Agent-local overrides remain intact; a conflicting shared-key replacement reports how to resolve the override.

## Evidence

- Baseline CLI: removing a configured profile failed with the provider-reference refusal.
- Baseline browser: Models reported “Secret saved,” while the key remained inline in provider configuration and the profile store stayed empty.
- Real Gateway/browser and CLI proof: save and removal produce identical credential/configuration state; both refresh auth successfully and preserve the selected model.
- Targeted owner, CLI, Gateway, UI, portability, and plugin-interface tests cover persistence, reference cleanup, concurrent credential replacement, retained credential kinds, explicit backups, administrator scope, and committed-write warnings.
- Independent real-browser acceptance passed UI/CLI state parity, metadata and credential-kind preservation, administrator-scope enforcement, config-write failure recovery, and active-run preservation during targeted removal. Full-provider logout closed the matching real local-provider streams.
- Final rebased candidate: 230 targeted owner, CLI, and Gateway tests passed. The rebuilt Gateway/UI passed the real-browser and CLI save/remove parity test. UI and portability checks passed on the integrated candidate, whose Auth B production files match the final candidate.
- Corrective focused checks: 80 tests passed for the binding race, public plugin boundary, CLI test types, and UI end-to-end inventory. Protocol generation passed.
- Real CLI race proof: a second CLI changed the provider from `fixture:manual` to `fixture:secondary` while key entry waited. The first CLI then preserved `fixture:secondary`, returned the saved-key recovery message, and exited 1.
- Final review-fix checks: 151 owner and Gateway tests passed. They cover configured and default reference-backed profiles plus API-key-only and full-provider removal races that preserve credential, profile metadata, order, and provider binding.
- Final public CLI proof: reference-backed replacement rejected without state change, then logout and save succeeded. A same-ID concurrent replacement completed while logout waited on the config lock; the rejected logout preserved credential and config, then retry and save succeeded.
- Rebased-head checks: 291 focused owner, CLI, Gateway, SDK, and UI-inventory tests plus 50 Models-page tests passed. Protocol generation, full build, the public CLI campaign, and real Models-page UI/CLI parity passed again after the Models login work landed.
- Incomplete-removal checks: 154 owner and Gateway tests passed. They cover thrown store errors, incomplete removal, partial multi-store deletion, surviving-reference restoration, and successful retry.
- Untargeted-binding checks: 155 owner and Gateway tests passed. Failed API-key-only removal preserves the retained token binding, and retry removes only the targeted key.
- Proof limit: the running Gateway's internal refresh exception was not live-injected. Gateway tests cover the failure branches and UI tests cover the resulting warning; remote-target and absent-local-Gateway CLI outcomes were observed separately.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-11 05:12:54 +05:30
Dallin Romney
458f9980c2
fix(release): set up Node before frozen admission (#144411)
* fix(release): set up Node before frozen admission

* fix(release): initialize Node for sibling admission planners

Extend the existing parent bootstrap to Release Checks, reusable source
admission, and the independently scheduled release matrix planner. Reuse the
trusted Node-only helper before each first TypeScript import while preserving
conditional parser installation and frozen-source admission checks.

Extend existing owner regressions and consolidate the parent-only assertion.
The stable release candidate remains unchanged.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-10 15:47:24 -07:00
RoboClaw
7444f28d2d
fix(ci): prepare selected release native fixtures (#144338)
* fix(ci): prepare selected release native fixtures

Enable and prepare the selected native heartbeat live test through its
existing runtime build owner. Keep Doctor scenario and canonical-path
service shims on the same selected checkout while retaining trusted shared
helpers. Complete the managed test's ephemeral TCP endpoint adapter so
shutdown verification cannot observe an unrelated host Gateway.

Qualification context: https://github.com/openclaw/openclaw/actions/runs/34507645027

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

* test(ci): bind live-shard cancellation to the test child

Use an isolated APNs-only inventory through the existing shard selector so
build preparation cannot masquerade as live-child readiness. Assert the
actual test:live arguments and join owned processes before fixture cleanup.

Preserve the original signal, descendant-death, and timeout assertions.
The original 20-file CI order passes all 332 tests after the correction.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-10 13:26:13 -07:00
Vincent Koc
cc9fe3963c
fix(ci): check frozen contracts before release fanout (#144236)
* feat(release): prepare npm and ClawHub for one-button publication

Stage complete plugin inventories in non-publishing owner workflows and
publish the original tarballs through the existing protected npm and
ClawHub publishers from one readiness receipt.

Bind source, tooling, producer attempts, and artifact digests. Require
established ClawHub publishers before readiness, verify every prepared
package before writers, and activate GitHub visibility only after exact
parent and canonical registry readback checks.

Share bounded artifact download recovery and verified archive reuse.
Retain partial dispatch requests, support explicit preparation adoption
and the existing successful-core resume path, and never blindly repeat
an uncertain registry mutation. Document native/platform boundaries and
the existing failed-core-child reconciliation limit.

Refs #136392

* fix(ci): run frozen bundle clients from their shipped layout

Resolve only the selected bundle contract with local committed-object reads and trusted syntax parsing. Preserve each legacy client's manifest, helper, import depth and bytes before Docker staging, and reject unknown or unreadable contracts.

* fix(ci): preserve frozen source read failures

Distinguish committed absence from missing, corrupt, or wrong-kind Git objects before selecting frozen compatibility. Share bounded local-only source reads across the shell and bundle resolver, and preserve errors through conditional callers before gateway Docker work. This is the source-read stage only; aggregate frozen admission remains separate.

* fix(ci): reconcile retained full release dispatches

* fix(ci): retain postpublish diagnostics when verification fails

* fix(ci): isolate frozen consumer contracts

* fix(ci): invoke publication diagnostics through guarded entrypoint

* fix(ci): add inert frozen target admission

* fix(ci): admit frozen source contracts before release work

Bind selected source, trusted tooling, package identity and resolved baseline selections at the four existing workflow prerequisites. Keep acquisition and conditional trusted parser provisioning separate from inert evaluation.

Share the actual consumer selectors, preserve authorized omissions and preparation-only assets, and retain bounded admission diagnostics without treating them as validation or publication authority.

* fix(ci): pin package tooling and align admission fixtures

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-11 03:11:58 +08:00
Peter Steinberger
bb3b67a5b0
fix(crabbox): reuse managed CLI across QA and remote proof (#144154)
* refactor(crabbox): centralize CLI discovery and setup

Reuse plugin-owned binary admission in QA Lab and workflow setup, carry verified versions through proof tooling, and preserve caller environments. Follow-up to #143768.

* fix(crabbox): preserve QA executable working directory
2026-09-10 09:57:28 -07:00