* fix(zalouser): avoid Gateway stalls during credential persistence
Move QR credential saves and logout revocation onto the shared SQLite worker,
await their durable completion, and keep restoration behind pending revocation
on its originally selected state directory. Preserve the credential format,
revocation markers, refresh CAS, and named older-host capability fallback.
Carry a live hosted-wizard assertion through setup and into the final write
grant so a detached QR login cannot persist after its owner is retired.
Validation: 172 focused tests across worker storage, registered setup, synthetic
SDK transport, and core wizard lifecycle; complete independent P2 review clean.
Repository gates: 34 passed, 3 unchanged core gates reused. The inherited unused
TSGO_CORE_TEST_MAX_ROOTS export is the sole remaining gate failure on this base;
canonical fix 1619aff478 supplies its repair.
Regression controls prove the former host-SQL logout path, a missing final
wizard-lifetime guard, and state-root drift across a pending logout. No live
provider account, operator Gateway restart, schema change, or protocol bump.
* test(gateway): await recap publication before projection assertion
Wait for the existing owner's next publication after releasing the model result,
assert its captured state without filtering successful outcomes, and check one
direct describe response. Preserve the durable watermark and model-call assertions.
No production logic or test deadline changes.
Validation: 22 cases on the exact tested CI merge with Node 24.21; real writer
queue probe preserves the old updating recap until write settlement, then exposes
the published current recap. Owning type graph, typed lint, formatting, and P2
review pass. The initial CI log redacts the returned summary, so the change does
not claim that its historical failure was proved timing-only.
* test(discord): scope policy reload fixture activation
Limit the synthetic policy fixture to its Discord plugin so real reloads do
not cold-load unrelated runtime plugins. Require the reload to report
applied while preserving the active-turn, held lookup and revocation flow.
Exact CI merge-source proof passed locally in 15.3s. Root fixture types,
typed lint, formatting and independent review passed. No production code
or test timeout changed.
* feat(browser): add opt-in Lightpanda semantic profiles
* fix(browser): reject direct selectors for Lightpanda profiles
* feat(browser): add portable Lightpanda deployment and benchmarks
* docs(browser): translate lightweight browser page title
* fix(browser): verify stale targets with a valid navigation control
* fix(agents): scope subagent concurrency to spawning sessions
Give each immediate spawning session its own configured execution budget,
including nested orchestrators, instead of making unrelated sessions share
one Gateway-wide queue. Preserve queue ordering, cancellation, hot limit
updates, and idle cleanup; retain the same parent identity during compaction.
Show aggregate activity with an explicit per-session limit in diagnostics.
Keep the existing child admission cap and Codex-native scheduling separate.
Verify with focused queue/runtime/UI tests and an isolated real-provider
Gateway case that holds one session's slot while another session and its
nested child run, then awaits all completion acknowledgments.
Closes#156370
* test(telegram): match sticker fixture to ingress dispatch
Register the native sticker pipeline's provider runtime prerequisite for
local and CI execution. Preserve the same Request JSON decoder while
exercising the registered bot handler, as durable webhook ingress does
after acknowledgment, and join the held turn before fixture cleanup.
The original CI file order reproduced the grammY timeout. Runtime
preparation alone was insufficient; the test's generic webhook adapter
imposed a deadline that is absent from the production dispatch boundary.
Keep the cache, media, and model-admission assertions unchanged, with the
real webhook acknowledgment contract covered by its existing test.
* test: repair packaging and runtime-owner CI fixtures
Copy the newly imported check-limits helper into the trusted packaging
harness so its dependency-free startup assertion reaches the CLI.
Compare prepared Telegram runtime files against the same config that
supplies the expected files. Keep worker-envelope coverage, packing,
concurrency and execution-budget assertions unchanged.
Both regressions reproduced before their fixes. The packaging owner passed
64 tests and ten targeted packing/prerequisite checks passed afterward.
* test(matrix): await follow-up adoption while active turn is held
Replace the 500 ms post-routing poll with resolver-owned adoption completion. Preserve foreground FIFO settlement by releasing the active turn before joining handlers, and settle fixture gates on timeout.
Validation: all four owner cases pass (37.31 s wrapper, 3.32 s bodies); selected changed checks and independent P0-P2 review pass.
* test(ci): recognize stopped Linux fixture process groups
Reuse the existing all-thread process-group assertion after detached fixture leaders are joined. Kernel kill probes also include exited zombies awaiting reaping; retain rejection of live descendants and uncertain observations.
Validation: 64 Docker scheduler cases and 18 retained-zombie/census cases passed on Linux Testbox; selected changed checks and independent P0-P2 review pass. The original CI process state did not reproduce in the original-order diagnostic.
* test(ui): count roster refreshes independently of child lookups
Install the browser clock before navigation and match roster sessions.list requests by includeGlobal. Preserve the 4999/+2 ms event window and avatar checks, and require exactly one roster refresh after the boundary.
Validation: reproduced the failing count and traced its extra request to a spawnedBy child lookup while the roster timer remained pending. All 14 browser cases pass (38.15 s; changed case 1.063 s). Selected changed gates and independent P0-P2 review pass.
Recognize plain, latest-alias, and dated Grok release IDs using xAI's documented capability ranges, so newer releases receive supported thinking levels and image input without another exact-ID update. Keep Grok 4.20 and variant suffixes on their existing conservative paths, and share the release rule with Responses tool defaults.
Validation: 56 focused regression tests passed on the refreshed head; independent review found no P0/P1 issues.
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
## What Problem This Solves
Intermittent tasks repeatedly recreate healthy Workers after the existing 30–60 second idle timeout, paying isolate startup costs and leaving resident memory behind after teardown. Existing metrics cannot distinguish which scripts are starting and exiting.
## User Impact
Hot task pools retain one Worker for five minutes of inactivity after an idle-retired Worker is promptly needed again. Node Code Mode retains its existing one-worker cache for five minutes. Operators can graph `openclaw_worker_started_total{script}` and `openclaw_worker_retired_total{script,reason}` through the existing diagnostics heartbeat.
## Why This Change Was Made
The existing retirement owner now owns idle timers and the single warm slot. Other slots keep their ordinary timeout; native cleanup, cancellation, pressure retirement, and rotation retain their existing ownership barriers. Existing idle GC collects released payloads in place. Code Mode still checks the runtime entry and heap limit before reuse.
The existing Worker resource registry owns cumulative lifecycle counts, including bundled direct Worker constructors. Counts advance on successful creation and confirmed native exit, survive exporter restart, and use bounded labels. Eval/unknown Workers remain `other`; nested Workers and V8-native threads remain outside the parent registry. Memory projection moves into a focused Prometheus module to keep the existing service below its growth ratchet.
No new configuration, database changes, dependencies, or managed-service environment changes. Updating installs the new runtime behavior through the normal restart; no migration or operator-state mutation is needed.
## Evidence
Real production-owner calls, ten verified synthetic tasks at 70-second simulated idle intervals (630 seconds observed), with real Worker creation and native teardown:
| Script | Starts before → after | Starts/min before → after | Task-run wall before → after |
| --- | ---: | ---: | ---: |
| `cron-stream-matcher.worker.js` | 10 → 2 | 0.952 → 0.190 | 485 → 111 ms |
| `git-operation.worker.js` | 10 → 2 | 0.952 → 0.190 | 3,076 → 824 ms |
| `code-mode-node.worker.js` | 10 → 1 | 0.952 → 0.095 | 858 → 187 ms |
The original source fails both new warm-window regressions; the pool creates a third Worker and Code Mode creates a second pool. In-place GC proof releases a 138 MiB payload to a 7.4 MiB heap while reusing the Worker.
### Allocator experiment and limits
Requested the largest advertised Blacksmith Linux class, 32 vCPUs, through a temporary proof-only workflow branch. The resulting guest exposes **8 online CPUs**, so this does **not** fulfill ≥32-vCPU validation or prove production memory behavior. Node 24.19.0, glibc 2.39, 32 concurrent Workers, 1,000 identical tasks touching 8 MiB each:
| Fresh process | Worker starts | Wall | RSS minus aggregate heapTotal increase after all Workers exit |
| --- | ---: | ---: | ---: |
| Create/terminate each task | 1,000 | 8.11 s | +639.59 MiB |
| Same churn, `MALLOC_ARENA_MAX=2` | 1,000 | 7.54 s | +423.86 MiB |
| Reuse 32 Workers | 32 | 3.08 s | +512.05 MiB |
The arena control reduces final residue 33.7%; reuse reduces it 19.9%. The churn proxy's post-100 slope is 102.63 KiB/task, versus 19.26 KiB/task with the arena control. A comparable retained-worker proxy slope is **not interpretable**: V8 heap capacity shrinks 992 MiB while RSS remains about 716.5 MiB, producing negative subtraction values and a false slope. Retained raw RSS grows only 0.152 MiB from tasks 400–1,000. These short exploratory results do not establish a many-core native-memory slope or allocator-contention tradeoff; consequently this PR does not set `MALLOC_ARENA_MAX`. Live deployment/production attribution remains a follow-up.
### Validation
Provider: `blacksmith-testbox`; profile: `openclaw-check`; lease: `tbx_01m36makaczyy9pds4yt5t2pem`; [workflow run](https://github.com/openclaw/openclaw/actions/runs/35834466664). One active remote command at a time, normal synchronization throughout. All test/typecheck/build work ran remotely.
- `node scripts/check-changed.mjs --base HEAD -- <all 25 changed paths>`: passed in 43m18s, including all 25 core-test typecheck graphs, core/extension production and test types, lint, documentation, SDK boundaries, dead exports, and import cycles.
- `pnpm test <file> --maxWorkers=1` for every changed test file, then `node scripts/run-vitest.mjs` for the five affected lifecycle/transport siblings: passed; 153 tests across 11 files. Final combined remote command: 59.2s. Original-source targeted runs failed both new regressions for the intended extra Worker/pool creation; candidate runs pass.
- The production-owner before/after benchmark command passed, including synthetic output assertions and native-exit cleanup for all three scripts. The allocator benchmark completed all three fresh-process cases with identical checksums.
- `pnpm build` followed by `OPENCLAW_LOCAL_CHECK=0 node --import tsx scripts/profile-extension-memory.mts --extension telegram --skip-combined --concurrency 1`: passed in 3m54s combined. All 95 plugin distributions built; bootstrap and 53 native control-plane module checks passed. Telegram isolated import exited cleanly, 160.09 MiB peak RSS (113.85 MiB above the profiler's empty-process baseline).
- Independent Codex review: scoped-clean, no actionable P0–P2 findings. `git diff --check`: passed. Net production growth: 146 lines; no test-only production seam.
Measured single-worker test command cost (seconds, including setup):
| Changed file | Wall |
| --- | ---: |
| `src/infra/worker-task-pool.test.ts` | 20.33 |
| `src/infra/worker-cpu.test.ts` | 7.48 |
| `src/logging/diagnostic-memory.test.ts` | 6.58 |
| `src/agents/code-mode-node.lifecycle.test.ts` | 5.21 |
| `extensions/diagnostics-prometheus/src/service.event-loop.test.ts` | 1.81 |
| `extensions/telegram/src/telegram-ingress-worker.test.ts` | 5.18 |
The two added behavior regressions use a fake idle clock; the real-worker pool case took 45ms and the Code Mode lifecycle case 35ms. CI timing will be updated after the head run exists. Crabbox reported an external runner-portal sync timeout after successful changed-check, final-test, and build commands; their underlying commands and reported run status succeeded.
Setup failures were diagnosed rather than counted as product failures: the transport checkout initially lacked `tsx` (fixed with remote frozen install), has no cgroup-v2 `cpu.max`, and lacks `origin/main` (changed checks use the actual original `HEAD` base plus exact changed paths). The initial benchmark fixture awaited an unrelated/unreferenced Worker exit; the corrected fixture uses the pool's resource-release boundary. The first changed gate identified line-cap growth; moving idle timing into its owner and extracting memory projection resolved it without an exception.
Close fixture responses so grammY cannot reuse a socket as the native HTTP server expires its idle lifetime. Preserve production ambiguity handling and the terminal-delivery assertions.
Refs #155040. Reproduced the PR #156189 run 35834470541 typing-only failure at the native idle boundary; the same case passes with the fixture header. Linux proof includes 30 delayed-response stress trials, 243 tests across 13 files, check:changed, and independent review.
Share concurrent Session Share node listings under service-owned authority and bound progressive chat-startup waits to five seconds. Retain compatible pending pages and publish completed refreshes through the existing catalog lifecycle, preserving current caller filtering. Targeted metadata, pagination, and non-progress clients retain complete-response semantics.
A synthetic six-caller progressive burst behind a 30-second node response improves from 30000 ms p99 and six RPCs to 5000 ms and one RPC. Testbox validation passed 41 focused tests, 64 existing gateway/service tests, typechecks, lint, and architecture checks. Independent review found no remaining actionable issues.
Carry operator role model limits through native Codex parents, child agents, restored work, reviews, and model-calling tools, with current-source revocation and lifecycle cleanup. Require an explicit Visitor model policy while preserving the documented staff-mode limitation and existing configured service authority.
The mock released child results when the parent's HTTP response was sent,
before the parent execution owner closed. Timestamp-prefixed all-settled
inputs also missed the fixture matcher and produced a generic response.
Correlate pending children with the exact runtime parent and use the existing
scenario wait loops to verify sessions.list reports that parent done, inactive
and not aborted. Keep the existing budgets and every direct-delivery, exact-send,
privacy and restart assertion. Recognize timestamped settlement inputs without
accepting quoted history. Migrate all mock-server and scenario callers together.
The six changed standalone test files each pass 20 full runs, with 91 focused
Gateway/QA cases and 43 HTTP sibling cases passing. Linux original profile 4
passes three times (27/27 scenarios, zero skipped), plus native Telegram and
QA-channel/empty flows. Actions proof: 35846151730. Single-worker file walls:
parser 2.47s, gate 2.48s, handoff 11.67s, routing 15.33s, surface 24.26s.
Types, scoped type-aware lint with a negative canary, formatting and P2 review
pass. Full lint declaration preparation hit ancestor-install isolation and used
the requested scoped substitute. Installed-package upgrade/rollback was not run
because no candidate tarball was configured; its migrated callbacks typecheck.
Closes#149640
## What Problem This Solves
Fixes: a stale processing progress message stays in Discord when a follow-up queued behind an active turn creates a draft after its inbound dispatch has returned.
## User Impact
The draft is cleared when the queued turn settles, the same as Telegram and Slack. No configuration changes or new plugin APIs are required. Failed, held, ambiguous, and partial queued finals remain a shared gap: their origin-route delivery outcomes are not available to channel settlement cleanup. Preserving drafts for those outcomes needs a separate core fix and is not claimed by this PR.
## Why This Change Was Made
Discord now uses the existing queued-follow-up settlement callback to run its draft cleanup, as Telegram and Slack already do. Cleanup belongs to the draft lifecycle rather than a new final-delivery hook in the public reply options.
### Maintainer decision
Ayaan's decision, September 23, 2026, as relayed by the coordinating maintainer:
> Land #156186 with parity to Telegram/Slack.
This accepts settlement-based cleanup for this Discord-only fix. The hosted finding about failed, held, ambiguous, and partial queued finals is acknowledged as a shared gap for a separate core fix, not dismissed as already solved. This PR changes only two Discord files; it changes no Slack files.
## Evidence
- The contributor's original late-queued-draft regression failed against the pre-fix production code.
- The regression adapted to the existing settlement callback also failed before the repair: the draft ID remained `preview-next` after settlement.
- Complete Discord draft-progress file: 30 tests pass (22.88 seconds). Draft-recovery file: 21 tests pass (20.23 seconds), including ordinary dispatcher delivered-error and failed-final retention. Those recovery tests do not establish retention for queued origin-route failures.
- Live Discord proof passed on `4857854a66de8824f88f87607fcaace265cbef3f`, using Convex-leased bots, an isolated Gateway, and the deterministic mock provider. The second tool turn was sent while the first progress draft was visible. Discord Gateway events showed its distinct progress draft, final reply, and draft deletion; REST history then confirmed the final remained and the draft was gone.
- The live doctor's readiness check passed. Owned fixture cleanup completed with zero failures and no unresolved operations. No manual Discord rendering or screenshot claim is made.
Redacted live event sequence (UTC, September 23):
| Time | Observed Discord event |
| --- | --- |
| 05:53:09.126 | First turn's progress draft created |
| 05:53:09.695 | Follow-up trigger sent while the first turn was still running |
| 05:53:18.758 | First turn's final reply created |
| 05:53:21.428 | Queued turn's distinct progress draft created |
| 05:53:25.485 | Queued turn's final reply created |
| 05:53:25.928 | Queued turn's progress draft deleted |
| 05:53:29.527 | Fixture teardown completed successfully |
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* refactor(channels): await operational account resolution
Prefer optional asynchronous account and configured-state hooks for runtime
operations while retaining synchronous SDK compatibility. Read Matrix
credentials through the worker-backed store and honor the supplied environment
when selecting its default account.
Preserve bootstrap metadata decisions and ordered target validation. Recheck
request, registry, and configuration authority around logout preparation and
teardown so reloads cannot redirect an older logout onto its successor.
* fix(channels): preserve logout results across async account changes
## What Problem This Solves
Fixes: new models (for example GPT-6 Sol and Luna from #155967) never reach existing installs because hosted catalog publishing has failed on every scheduled run since 2026-09-21, after models.dev stopped publishing the `kimi-for-coding` provider.
## User Impact
User impact: once this lands and the catalog republishes, installs on older builds get current model metadata on their next Gateway restart, as intended, without a release. A renamed or broken upstream provider now degrades only that provider's models.dev hydration instead of freezing catalog updates for all 44 providers. No config, schema, or runtime change.
## Why This Change Was Made
- The Kimi plugin mapped `kimi` to models.dev `kimi-for-coding`, which no longer exists. models.dev split it into `kimi-code-plan-global` (api.kimi.ai) and `kimi-code-plan-cn` (api.kimi.com). OpenClaw's provider uses `https://api.kimi.com/coding/`, so the mapping now points at `kimi-code-plan-cn`. Its four model ids match the manifest.
- The publisher threw for any missing or malformed mapped upstream provider, which failed the whole publication. It now logs a warning naming that provider and publishes its manifest rows unhydrated. A models.dev outage or a non-object response stays fatal and still preserves the last published artifact.
- Docs describing hydration failures are updated to the new contract.
## Evidence
- Production failure: `openclaw/catalog` "Publish catalog" runs on 09-22 and 09-23 end with `models.dev catalog missing or malformed for provider kimi-for-coding` / `[publish-model-catalog] FAILED (exit 1)`. The published `models/v1/catalog.json` was last changed on 09-18 and has no `gpt-6-sol`.
- Downstream symptom: a Gateway built before #155967 kept a synthesized placeholder for configured `openai/gpt-6-sol` (128K, text-only, no reasoning) after restart, so its thinking level clamped to `off`.
- Live dry run of the workflow command (`publish-model-catalog.mts --pricing`) against current models.dev and pricing feeds:
- `origin/main` (`e7fbe5389e0`): reproduces `missing or malformed for provider kimi-for-coding` / `FAILED (exit 1)`.
- This branch: exit 0, `providers=44 models=1062`, `models.dev provider=kimi added=0 filled=0 skipped=0`. The written bundle contains `openai/gpt-6-sol` and `gpt-6-luna` with `reasoning: true`, text+image input, 1,050,000 context window and 272,000 context tokens. The 10 remaining warnings are pre-existing pricing gaps.
- End to end on a Gateway image built before #155967 (fresh state, no network, catalog served locally): `openclaw models refresh` against the currently published catalog reports `unchanged (44 providers, 1033 models; generated 2026-09-18)` and has no `openai/gpt-6-sol` row. Against this branch's published output it reports `updated (44 providers, 1062 models)`, and `models list` returns `openai/gpt-6-sol` as text+image with a 1,050,000 context window and 272,000 context tokens.
- Tests: `test/scripts/publish-model-catalog.test.ts` 73/73 pass. The new case fails on the original script with the production error message, and passes with the fix. Outage and malformed-feed cases still reject without changing the bundle.
- Test cost: `pnpm test test/scripts/publish-model-catalog.test.ts --maxWorkers=1` took 21.1 s wall on an Apple M3 Ultra (Vitest duration 18.61 s, 73 tests; 62% transform, 34% tests). CI run 35837009655 on this head passed all 117 jobs; its logs report no per-file timing for this test.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Project bounded Desktop metadata before cache and overlay retention, preserving the independent CLI index contract. Detach trimmed strings so short display values cannot keep large backing strings alive.
Synthetic full catalog proof reduced additional retained heap from 64.49 MiB to 0.482 MiB while preserving all 48 visible rows and metadata. Fixes#155754.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* refactor: retire pre-June import and verification compatibility
Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.
Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.
Refs #156190
* docs: route legacy upgrades through 2026.9.5
* test: await Telegram fixture lifecycle events
Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
* feat: extend Ultra harness mode across supported runtimes
* fix: preserve Ultra across model capability boundaries
* test: align remaining session fixtures with Ultra harness mode
* test: bind Ultra session fixture to prepared provider policy
* test: align Ultra capability fixtures
Keep the full Gateway assertions while isolating observed native efforts from known model capability floors. Use an unsupported native effort in the compaction clamp fixture now that Ultra is a supported harness mode.
* fix: honor configured Ultra capabilities across runtimes
Compose configured model overrides with captured catalog rows for CLI and cloud execution in omitted, merge, and replace modes. Keep route-bound capabilities with their transport owner, including when applying a previously unspecified route.
Consolidate Telegram tests at their behavioral owners and remove redundant mock inventories and private test seams while retaining distinct routing, authorization, replay, media, and streaming regression checks.
Align retained delivery tests with the landed shared final-delivery lifecycle. Respect replyToMode off when finalizing a quoted preview so the answer is edited in place instead of deleted and re-sent.
Exact-head hosted CI passed. The PR records current-head Telegram Test Server proof with reverted controls and the remaining validation limitations.
Work session: https://team.openclaw.ai/chat/roboclaw/dashboard/7320cd5b-9414-4cf6-99fe-0abb9a51df58
Co-authored-by: obviyus <22031114+obviyus@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Feed captured codesign metadata directly to grep so successful matches cannot leave printf failing with SIGPIPE under pipefail. Extend the existing desktop fixture with a bounded diagnostic tail that exposes the race deterministically.
Carry the prepared workspace worker cut onto current node launch, workspace,
and Codex lifecycle owners. Preserve synchronous external plugin acquisition,
transaction-time authority, and durable retirement semantics.
Join invocation and duplex cancellation through the runtime handoff and fence
legacy synchronous acquisition while async workspace mutation admission waits.
Private source checkpoint: fresh scoped review is clean through P2 and 43
selected ordinary/regression cases pass. Full changed checks and final build
remain incomplete; assertion baseline shrink maintenance and the current-main
schema 18 carry follow before qualification. The historical specific
reproduction remains held and was not replayed.
Fixes#156008
Related: #127605 (Mattermost slice)
## What Problem This Solves
A reaction on a Mattermost channel-thread message reaches the parent channel session instead of the thread containing that message.
## User Impact
Thread reactions now reach the same session as their messages. Top-level reactions remain in the parent session with default threading off. Configured threading follows existing message routing; failed post lookups retain the parent-session fallback. No configuration, public API, or stored-data changes are required.
## Why This Change Was Made
Reaction payloads contain a post ID but no thread root. Resolve the post through the existing bounded resource-cache pattern, then let the existing event plan choose the session using the same post identity as the message path.
## Evidence
- Exercised the real REST client, monitor resource cache, reaction handler, and thread resolver against a loopback Mattermost HTTP stub. The runtime event sink records the selected session.
- With pre-fix production sources, the thread assertion fails: the observed key is `mattermost:default:channel:chan-1`. With this head, it is `mattermost:default:channel:chan-1:thread:root-1` and the server observes the post lookup.
- Top-level/default-off and unresolved-post controls retain the parent session on both sides.
- Candidate regression and resource suites: 31 tests passed. The supplemental loopback proof adds three passing scenarios.
- Simplification retained the existing routing owner and bounded metadata-cache pattern; no competing policy or public seam. Production delta: +41/-2; retained test delta: +196/-0.
- Exact-head hosted CI run 35762478168 is green; hosted ClawSweeper is ready for `2f2942660abd633dcd4775d03b7dff7d4c547f43`.
- No live Mattermost tenant or WebSocket transport was exercised. Documentation already describes routed reaction events and existing thread settings; no documentation or release-owned changelog change was needed.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Related: #155190 (provider-startup portion only; telephone-call teardown remains separate).
## What Problem This Solves
Fixes GPT-Live sessions failing to start when a recoverable provider error arrives before `session.started`.
## User Impact
Recoverable warnings no longer prevent a subsequent readiness event from starting the voice session. Fatal authentication errors still reject startup immediately. No configuration or migration changes are needed.
## Why This Change Was Made
The bridge now checks the existing fatal-auth classification before rejecting startup, logs a redacted warning for recoverable startup errors, and continues waiting within the existing readiness timeout. Established-session error behavior and redaction remain unchanged. The classifier is unchanged: status 401 (number or string), or codes `authentication_error`, `invalid_api_key`, `invalid_token`, and `token_expired` are fatal; other error events, including `missing_scope` without status 401, are nonfatal.
Provider-error dispatch stays in a small private sibling module because the bridge is already at its enforced line cap. The regression checks readiness and terminal behavior rather than exact warning wording.
## Evidence
- Real bridge `connect()` replay using a real `ws` client and a TCP loopback WebSocket server, production event parsing/lifecycle, and the existing media-runtime test adapter. The server sent `error` followed by `session.started` after receiving `session.start`.
- Baseline `4d50b52bce`: `missing_scope` rejected startup; connected=false, ready callbacks=0. Candidate: startup resolved; connected=true, ready callbacks=1, warning=1, error callbacks=0, close callbacks=0.
- `authentication_error` rejected startup on both baseline and candidate, with connected=false and ready callbacks=0 despite the subsequent readiness frame.
- `node scripts/run-vitest.mjs extensions/openai/realtime-quicksilver-bridge.test.ts --maxWorkers=1`: 37 passed; measured command wall time 38.71 seconds (Vitest 34.73 seconds). Earlier exact contributor-head CI changed-extension test-shard step took 133 seconds.
- Replay ran under Bun 1.4.2; this is synthetic provider-stream transport proof, not live OpenAI availability, full Gateway, telephone-call, or media-worker packaging proof. No real provider credentials were used.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Grok (xAI) Talk saved each spoken sentence as several duplicate or truncated user messages, because every cumulative transcription snapshot was persisted as a final turn. Longer spoken replies were also cancelled when they overflowed the browser's 10 s playback queue. Tool-call consults did not tell the agent which blocked call the user had confirmed.
The xAI provider now previews input snapshots and commits one final per utterance at a real boundary: next speech, response end, session close, or 1.5 s of quiet after a late recognition. It fences output, cancel errors and buffered tool calls from retired responses. The relay forwards snapshot mode and the saved transcript id, so the Control UI hides a live caption once its saved row arrives. Browser playback allows 60 s / 4,096 sources, and the 20 ms relay frame contract is unchanged. Tool-call consults receive the same blocked-call retry context and confirmation-id reply as native delegation. A same-action retry in a new run reuses the pending challenge without extending it. `onTranscript` gains an optional `{ textMode: "snapshot" }` metadata argument.
Proof: live isolated Gateway and Control UI Talk runs on xAI Grok against main.
- Three utterances were saved as 3 turns instead of 22. Long answers played to the barge-in instead of being cut by overflow.
- A spoken "Yes." wrote the confirmed file exactly once; "No." wrote nothing.
- Non-affirmations, expired challenges, superseded or consumed ids, and ids from a closed session were all rejected before the exec ran.
Co-authored-by: Marvinthebored <peter@lindsey.jp>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(agentsapi): include conversation context in native messages
* refactor(agentsapi): use private attempt context helper
* refactor(agents): keep inbound context type with shared owner
* fix(gateway): preserve context for active harness messages
---------
Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>