Two smoke guides still described the hidden Node installation retry removed by #161512. Describe the single install with the existing Bun launcher marker, its terminal failure behavior, and Node's preparation-only role.
Both guides pass formatting and whitespace checks; independent review found no actionable issues. Hosted preflight failed in unchanged main CI because it invoked the absent .ci-harness manifest path. That failure was reproduced on detached main and qualified through the native pre-existing-failure admin exception; security-fast passed independently.
* fix(test): include omitted agent tests in directory and glob runs
Expand agent directory and ordinary glob selections into existing regular
files before choosing their runners. Preserve fast, isolated, harness and
database-worker ownership, exclusions, include limits and CLI controls.
Replace plan-shape assertions with actual configuration inventory proof.
The regression fails on the original selector with 163 of 191 tools files
selected. Keep complete coverage while avoiding quadratic test matching.
Validation: Linux Testbox owner and sibling suites, real directory/glob
commands, build, changed checks, import boundaries and both cycle scans;
independent review through P2 passed.
* test: batch helper-import coverage parsing
* refactor(infra): share filesystem observation policy
Centralize polling overrides, guarded source admission and metadata sampling. Bound remote file notifications and join accepted output before shutdown.
* refactor: migrate config, skills, memory and dev watchers
Use fs-safe invalidations and scope replacement while keeping source selection, settling, reload and indexing policy with each owner. Remove first-party Chokidar, bespoke native transports and duplicate observation tests. Preserve joined application work and watch-limit degradation; isolate manual Gateway writer fixtures from filesystem readiness.
Supersedes #158182. Thanks @vincentkoc.
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(memory): use host filesystem observation for captured plugins
Pair watch with SDK Root admission so captured plugin dependencies cannot split
fs-safe's module-local Root registry. Preserve the strong Root check and route
existing watcher tests through the host SDK boundary.
Keep the eager Skills subscriber snapshot with Array.from and remove the new
lint suppression without changing the production suppression allowlist.
* fix: preserve fs-safe fallback for legacy polling overrides
Treat legacy false, zero and empty polling overrides as a native preference
so Bun and installations without the native addon retain working observation.
Capture Config polling-recovery eligibility per observer to preserve the
explicit override retry limit.
Prepare the Skills recovery test through its existing worker owner and load
the real snapshot dependency before cases begin. Keep recovery assertions,
timeouts and artifact verification unchanged.
* test(memory): check resolved observation backend health
* fix(plugins): preserve native capture authority and lifetime
* test(ci): scope native Doctor proof to its runtime owners
* fix(plugins): break capture storage type import cycle
* fix(memory): import observation types through host SDK
* test(gateway): isolate operation journal fixtures
* fix(qa): point worktree lifecycle scenario at surviving run-end cleanup tests
#160308 deleted src/agents/worktrees/service.run-end-cleanup.test.ts, which the
managed-worktrees-workboard-lifecycle scenario still listed as a codeRef, so
extensions/qa-lab/src/scenario-catalog.test.ts failed on main. The surviving
run-end cleanup outcome coverage lives in service.test.ts (late claims, stale
lifecycle writes) and service.removal-safety.test.ts (dirty retention).
(cherry picked from commit 3a300c650a)
* test: isolate planner contracts and share installer shell
* refactor(gateway): schedule remote skill refresh (#160318)
## What Problem This Solves
Remote-node skill refresh still kept a private debounce timer and threaded its handle through Gateway startup and shutdown.
## User Impact
Skill changes keep the existing 30-second debounce, now owned by the kernel scheduler. Shutdown joins an active refresh and suppresses late broadcasts. Updating needs no operator action or config, storage, or public SDK migration.
## Why This Change Was Made
Schedule refresh directly with the existing scheduler. Remove the timer getter/setter, delay option, runtime handle, and redundant close hook.
Production **+18/-36/net -18**; tests **+100/-38/net +62**; docs **0**. Production counts use src/** and extensions/** with the work-order test exclusions.
## Evidence
Deslop and Codex autoreview completed with no actionable findings through P2. Focused proof passed on blacksmith-testbox lease `tbx_01m3k4y56rf9r63tzs4hck3pek`, [run 36378613468](https://github.com/openclaw/openclaw/actions/runs/36378613468); unchanged source/test hashes were verified in the later proof. Each command used `pnpm test <file> --maxWorkers=1`:
| File | Tests | Command wall |
|---|---:|---:|
| src/gateway/server-startup-early.test.ts | 14 | 14.763 s |
| src/gateway/server-close.test.ts | 80 | 23.385 s |
| src/gateway/server-startup-lifetime.test.ts | 14 | 37.299 s |
The closeout docs PR will carry the complete census and merged-main sleep/shutdown proof; that live proof is still pending.
Exact-head `node scripts/check-changed.mjs --base 5362ba0ca3 -- <all 6 changed paths>` passed on `4fafc2add7`: production typecheck, 25 dependent test type graphs, lint (0 warnings/errors), dead-export scans and boundary guards. Check wall **1610.89 s**. Provider blacksmith-testbox, lease `tbx_01m3khvbn4fa9jb0bhe5qqp3we`, [run 36397188865](https://github.com/openclaw/openclaw/actions/runs/36397188865). The candidate was materialized in a clean detached worktree because native sync retained its hydration HEAD.
## Inherited CI failures and landing evidence
Completed exact-head CI [36402003337](https://github.com/openclaw/openclaw/actions/runs/36402003337) has two underlying failures plus its aggregate gate. Both match independent PRs:
- `gateway-agent-skill-refresh.e2e.test.ts:304` (called from line192): expected lifecycle count3, actual2. Identical without this cutover on approval [run36401825740/job108861794753](https://github.com/openclaw/openclaw/actions/runs/36401825740/job/108861794753). The unchanged synchronous `src/skills/runtime/refresh-state.ts` producer invokes the test's earlier registered listener directly. This PR changes a separate downstream remote-bin refresh consumer; the failing lifecycle count precedes the debounce/broadcast assertion.
- `subagent-completion-blocked.e2e.test.ts:78`, ordinary delivery exhaustion: expected suspended, actual pending. Identical on questions [run36403551682/job108867383832](https://github.com/openclaw/openclaw/actions/runs/36403551682/job/108867383832). Neither PR touches that completion owner.
Focused tests and the complete exact-head Testbox gate passed. The maintainer's standing instruction authorizes pinned admin squash over these inherited failures. No CI rerun or weakened assertion was used. The installed wrapper rejects the CLI admin argument shape; the native workflow's protected GraphQL merge route preserves the same squash, message, and expected-head payload used by that CLI operation.
(cherry picked from commit df785c0971)
* test(gateway): adapt skills proof to fs-safe lifecycle
* test(skills): normalize Windows observation lookup
* test(memory): isolate guarded reconciliation from native hints
(cherry picked from commit f07aececc0b5641a00adfd0271c89e7c7162d38d)
* test(channels): update Synology context builder inventory
Match the selected builder spelling after the Synology inbound route was inlined in #160639. Preserve the existing caller inventory assertion and production behavior.
* test(skills): assert subscription renewal during recovery
* fix(ci): exchange hybrid parallel groups within the job cap
Reuse the existing bounded exchange optimizer for final hybrid consolidation. Preserve group ownership, child worker limits and admission budgets while reducing the fs-safe composition from 80 to 78 rows. Scope exchange by backend so valid Blacksmith layouts and cap-independent plan identity stay stable.
* fix(test): retire subagent sweepers before SQLite owners
Stop the original subagent registry before non-isolated file cleanup retires
its SQLite owners. Join accepted sweeps and cleanup tails while preserving
synchronous fixture resets, successor scheduling, and the actor path guard.
Await the new reset contract in the concurrency benchmark too.
Baseline controls demonstrate an old registry tick recreating a retired
shared-state owner and premature settlement of three in-flight retirements.
The receiving candidate passes 102 owning/sibling cases, including the
17-case nested fixture (16 passed, one expected skip). The benchmark with
eight child sessions also passes. The owning run took 156.42 seconds.
Types, lint, formatting, export scans, and repository guards were qualified.
The exact historical hosted timer schedule remains unobserved.
* test(ci): isolate worker boundary fixtures from prebuilt mode
Keep synthetic historical runner fixtures in their explicit no-dist mode so
an inherited CI prebuilt flag cannot enter package preparation before the
worker-owner boundary. Preserve all six fixture modes and their assertions.
* test(ci): align runner fixture with canonical main
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Add API and Claude CLI catalog entries with the published 1M/128K limits and
Sonnet 5 pricing, roll the bare `sonnet` alias to Sonnet 5.5 (`sonnet-5` stays
pinned), and give the model its own contract: `/think off` sends Anthropic's
`between_tools` setting instead of the rejected `disabled`, forced tool choices
relax to `auto`, retained thinking is prefix-bound with append-only runtime
context, a session moving from Sonnet 5 or Claude 4.x onto Sonnet 5.5 keeps its
reasoning while Opus 5, Fable, and Mythos reasoning is dropped, and direct
API-key requests opt into the safety refusal fallback to Sonnet 5. Fable 5.1 no
longer replays Sonnet 5.5 thinking, matching Anthropic's live replay behavior.
* fix(file-transfer): declare node commands idle between invokes
A default headless node host could never pass the auto-update idle
barrier: file.stat, file.create, workspace.memory and workspace.skills
registered without hasActiveWork, and the node host treats a missing
idle hook as active work. Their work is invoke-scoped and already
counted by the runtime's in-flight invoke tracking.
* test(e2e): prove default-plugin nodes activate node auto-updates
* test(e2e): read the node-host subsystem console prefix
* docs(install): explain recovering 2026.9.6 nodes stuck before automatic updates
* fix(gateway): admit shared-venue bursts and explain busy retries
* fix(ui): keep liveness probes out of the service worker cache
* refactor(ui): keep browser transport errors with the socket adapter
* test(ui): model refused socket state transitions
* style(ui): satisfy retry-path lint rules
* chore(ui): retire the removed gateway type assertion allowance
* fix(test): SQLite teardown fails when a test's agent database close overlaps the next test
Most fixtures tear down with the synchronous closeOpenClawAgentDatabasesForTest(),
which only schedules asynchronous agent-database Worker retirement. A Testbox
sweep of the 308 candidate files that never await the async drain found ~36
whose teardown leaves those closes running into the next test. A reclamation
lease release that overlaps later work and ends unsettled falls back to
cleanup that needs the host SQLite broker, which a Vitest thread never has,
so the runner fails the file and retains SQLite custody for the rest of the
worker (CI run 36290393694).
The shared runner now waits for scheduled agent database closes before each
test (ahead of aroundEach setup), before each retry attempt, and before a
file's cleanup phases. Failure reporting is unchanged: a failed close stays
with its owner and the file-end drain. Child-Vitest fixtures prove the next
test starts only after a scheduled close settled, including after a runtime
skip and around aroundEach setup and teardown.
* test: settle close-ordering fixtures on an event-loop turn
The ordering fixtures used a fixed 100ms timer, which the test policy
disallows and which could let the old runner pass if the inter-test gap
grew. Settle each synthetic close one event-loop turn later instead; the
teardown-to-next-test path is promise-only, so the base runner still
reaches the next test with the close pending (3 of 3 base runs fail,
3 of 3 patched runs pass on Testbox).
Worker-backed state writes leave a worker holding an idle state-database
connection. The synchronous closeOpenClawStateDatabaseForTest() only starts
that worker's retirement, whose last-connection checkpoint and WAL delete then
land at an arbitrary later time. Fixtures that opened the same database raw or
snapshotted/copied its files right after could hit SQLITE_BUSY or see bytes
change underneath them.
Await closeOpenClawStateDatabaseAsync() at every site a Testbox custody probe
showed exposed (same database, live worker custody), make
seedLegacyV15ProposalRows() async for the same reason, pin the owner contract
in the database-worker WAL test, and document the rule for fixture authors.
Fireworks retired accounts/fireworks/routers/glm-5p2-fast. Onboarding, the bundled catalog and the curated live selector now use the documented GLM 5.3 Fast router (Fast cached-input price 0.39/M); explicit operator pins of the old id are preserved.
Forward-port of release/2026.9.7 066fd58ec8.
Add an advisory Bun-only runtime smoke to Install Smoke, run only by Release Checks (Full Release Validation). It
installs the verified candidate with the pinned Bun fork, hides every real Node inside a private mount namespace,
and exercises install, CLI, Gateway, node host pairing, a mocked agent turn, doctor, terminals, and the browser.
A PATH sentinel records every node/npm/pnpm/yarn/corepack execution with its exact path and process ancestry, a Bun
preload records the JS stack, and a classifier fails on any attempt missing from
scripts/e2e/lib/bun-only-runtime/expected-node-blockers.json or on a listed blocker that no longer reproduces.
Admin-merged with maintainer approval: the only failing required check was test/scripts/pr-wrappers.test.ts, a main
regression from b2e0e562 unrelated to this change. PR-owned tests (94), lint, workflow checks, check-dependencies,
docs, guards, and types passed at this head.
Resolve the probe route from the selected scenario attempt's surviving
messages and use the provisioned primary participant. Keep the first
sample fresh and reject replies from a different topic.
Preserve explicit probes and native observer admission. Rework the
original repair without migrating the scenario catalog.
Related: https://github.com/openclaw/openclaw/pull/137139
OpenAI shut down its Sora video API: sora-2 and sora-2-pro report
shutdown_date 2026-09-24 and POST/GET /v1/videos now return HTTP 404 for
every project. Every video_generate call routed to openai/* failed with an
opaque "OpenAI video generation failed (HTTP 404)", and because the bundled
openai plugin still advertised a configured video provider, agents on
OpenAI-only installs kept selecting it.
Remove the OpenAI video-generation provider, its manifest contract and
metadata, live-test defaults and workflow filters, and the docs that
advertised Sora. Stale openai/sora-* refs need no migration: the media
runtime already skips them with "No video-generation provider registered
for openai" and continues to configured fallbacks or auto-detected
providers.
Remove Tasks and TaskFlow runtime, APIs, CLI, SDK surfaces and panels after the Cron, session, native execution and media completion ownership cutovers. Preserve stored rows and import provable legacy native assignments through Doctor; ambiguous ownership stays untouched with a warning.
Follows #158221, #158217, #158225, #158222, #158702 and #158776. Related: #156532. Task-specific public APIs retire immediately; retained responsibilities use their existing owners.
Maintainer-authorized administrative landing after full CI run 36312986498 attempt 2 passed on 274595e2, with subsequent actual conflicts reviewed and focused checks passing. Current PR CI preflight hits the 64 KiB changed-path metadata limit before tests (run 36335042695); its duplicate security-review status mirrors that planning failure. Review and scoped proof are recorded in the PR. Published 9.4 native import is proven; remaining native completion and 9.4 rollback witnesses are explicitly unproven.
Use transactions on the actual SQLite state and device databases for ordinary writes. Remove redundant coordination databases, transport, and exclusion layers while preserving bounded process ownership for startup, schema work, and offline maintenance.
Tie test and QA scratch retirement to settled workers and native resources, preserve active plugin captures, and join SDK declaration compiler processes before synchronous semantic rendering.
Validation: main-tier CI on 5e731c1f64 had 144 successful jobs and one Windows ACP initialization timeout. Qualified unchanged replay 36314027585 passed all 896 tests with the original 48-file order, six projects, toolchain, and deadlines. The original timeout remains unexplained and recorded in the PR. Reviewed main-conflict integration through 39caa592ef passes focused SQLite, Doctor, image, and Cron proof plus affected typechecks and lint. No accepted actionable independent-review findings remain.
Squash landing of #157413 under explicit maintainer authority to resolve logical main drift and admin-merge using the completed CI evidence. No PR-specific schema or public configuration migration.
* fix(agents): keep spawned workers out of the current conversation
Agent-spawned workers (sessions_spawn thread=true, native and ACP) may only bind a new child thread. Channels whose spawn placement is the current conversation (Telegram, Feishu, LINE, generic current-conversation bindings) now reject thread=true instead of handing the user's chat to the worker. Legacy spawn-created bindings on those channels are ignored by the binding service, so the conversation routes to its normal agent again. Discord/Matrix child-thread sessions and defaultSpawnContext are unchanged.
* test(telegram): expect parent routing after worker-takeover upgrade
* test(agents): refresh prompt snapshots for child-only thread spawns
* test(telegram): keep parent-phase checkpoint codes in public upgrade evidence
Use Podman native init for PID 1 and verify HostConfig.Init before starting the isolated test container. Native diagnosis found the terminated watchdog retained as a zombie with the same process-start identity; the unchanged disappearance assertion passes with init. Preserve all isolation, cancellation, joining, snapshot-retention and assertion budgets; document the existing init-helper prerequisite.
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
Scope the missing-load-path fixture to published drivers that admit invalid configuration before staging. Retain ordinary update and migration checks for older drivers, preserve exact-row reruns, and record fixture applicability in the published receipt.
* fix: preserve owed final replies when restart retires the sender
* fix: retain live owner checks during restart delivery
* refactor: consolidate final delivery rejection handling
* fix: distinguish source read failures from owner revocation
* fix: recover queued finals with original owner authority
* fix(agents): reject finals after source read failures
Keep custody retryable only for recognized restart retirement. Unknown source assertion failures cannot transfer continuity checks to recovery after the live closure disappears.
Prove source-read failure followed by intervening input cannot replay at either adapter handoff, with bound and unbound owners. Preserve genuine restart retention and revocation coverage; retain suppression reason literals in the batch recovery table.
* fix(ci): register channel owner policy scenario with Knip
The upgrade-survivor shell invokes this CLI by path. Model it in the shared executable-root list alongside sibling scenarios so the full-tree unused-file audit recognizes the real consumer.
The canonical unused-file check reproduced one unused file before registration and passes with zero entries afterward. pnpm deadcode:full and formatting also pass.
* test(gateway): await watcher-owned notice completion
* refactor(telegram): move bundled thread bindings to workers
Await binding hydration and persistence through the existing plugin-state worker. Keep public synchronous SDK compatibility on the same owner while bundled callers use awaited operations. Preserve FIFO admission, current authority, committed acknowledgements, and shutdown drainage.
Validation on the frozen source: Focused binding, bot lifecycle and ingress tests, extension production and test typechecks, scoped lint, fresh review, a normal full build, and the built Telegram import profile.
This private checkpoint retains proof from base 78d5bedb34. It does not resolve the inherited TS1619/current-main qualification limit or the unchanged baseline TSGO_CORE_TEST_MAX_ROOTS unused-export finding. ACP startup metadata reads remain separately owned.
* test(telegram): dispatch sticker checks through the admitted handler
* test(telegram): verify published-driver binding upgrades
* test(telegram): keep upgrade runner private
* fix(telegram): await binding activity on bundled routes
* test(telegram): bind stop fixture to the active routing facade
Missing prerequisites for two test lanes, a copy-paste slip in an SDK
sample, three undeclared identifiers in a quick start, a colon promising
keys named two paragraphs later, duplicated Related lists, two unlinked
pages that exist, an unbracketed placeholder, and a bare doctor invocation.
The thirteenth fix, an ffmpeg prerequisite on docs/tools/tts/quickstart.md,
is held back: every PR touching that page is refused by the secret scanner
before review.
Share build-scoped compile-cache ownership between the launcher and runtime. Reuse inherited namespaces, avoid redundant respawns, and retire superseded builds with best-effort seven-day and 512 MiB maintenance.
Testbox: 304 focused tests passed with one Bun-only skip; pinned changed-file checks passed. Thirty child launches used 41,544 bytes instead of 1,185,000 bytes, and same-build launch time fell from 2.46 s to 1.22 s. Persistent external-plugin source capture reuse remains outside this scoped change.
* ci: narrow PR tests by import and dependency impact
Resolve affected tests through the existing runtime import graph, workspace
and SDK aliases, configured Vitest inputs, and explicit policy owners.
Compare exact-base dependency closures instead of treating every lockfile
or package metadata change as global. Keep documentation and localization
data out of PR Node test rows.
Preserve canonical execution metadata and complete automatic main coverage.
Retain diagnostic fallbacks for global inputs and unverifiable ownership.
The 40-run replay reduces explicit fallbacks from 30 to 6 and median selected
jobs from 132 to 122, with no unexplained historical failure omissions.
The remaining large import closures do not meet the requested cost target.
* fix(ci): run test selector parsing under Node
Bun workers cannot provide the Node TypeScript parser used by the selector. Resolve the existing Node executable at the scanner boundary and retain that built-in-only dependency in materialized wrappers and standalone fixtures. Preserve real shard inventory exports in the command planner fixture.
* test(ci): isolate changed-gate compiler fixtures
Source-aware selection reaches inline compiler checks that bypassed the fixtures’ subprocess stubs, repeatedly compiling the real repository during gate-order and root-lint tests. Bind that owner to the shared synthetic recorder while preserving the real CLI, lint execution, serial-order assertions, and failure sentinels.
* fix(ci): retain compiler gates for TypeScript catalogs
Catalog-only PRs suppress Node test rows, but their TypeScript modules still need compiler validation. Admit the existing check owner independently of Node tests while preserving documentation and JSON-only skips. The catalog fixtures retain main and manual coverage and distinguish compiler admission from test execution.
* refactor: retire pre-June config and upgrade test support
Retire obsolete Doctor keys and their runtime fallbacks. Older configurations use the documented 2026.9.5 Doctor bridge; current upgrade verification starts at June 2026 while historical receipts remain readable.
* fix: preserve retired config when the upgrade bridge is skipped
* test: align upgrade fixtures with retained migration contracts
The project runner replaced timed-out configurations with another attempt
and could turn their failed hooks into a green job. CI run 35928401414,
job 107409413294, masked seven settlement-hook failures this way.
Remove replacement dispatch and its expanded retry deadline. Keep the first
watchdog failure even when a child exits zero, retain incomplete reports,
and preserve process/cache cleanup.
Proof: the temporary fail-first test failed the real CI shard entry point
with exit 143 and one attempt. The runner owner passed 20 standalone runs
and three CI-config replays on Testbox; shared-helper importers passed.
* ci: shorten real-Gateway UI validation
Use runtime-only preparation while build-artifacts retains SDK declaration checks. Preserve four exhaustive Gateway tours in full manual and release validation with their faster owner-boundary siblings in ordinary CI.
* test(ci): align real-Gateway preparation guards
Keep the paired runtime and UI build contract, assert the existing runtime-only mode, and document SDK validation ownership in the artifact job.
* ci: split real-Gateway UI validation into two jobs
* fix(ci): declare the shared real-Gateway parallel inventory
* test(ci): require manifest strings before decoding
* ci: balance standalone UI proofs across Gateway shards
* ci: defer desktop transport tour to release validation
* perf(test): reuse prebuilt Control UI assets
* test(ui): align Gateway fixtures with catalog and build contracts
* refactor: retire pre-June import and verification compatibility
Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.
Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.
Refs #156190
* docs: route legacy upgrades through 2026.9.5
* test: await Telegram fixture lifecycle events
Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
Use process workers on Windows so concurrent Vitest workers do not share temporary inheritable subprocess pipe handles. Preserve worker limits, file parallelism, and strict process/output cleanup. Route shared worker-policy changes through Windows CI.
* fix: keep web chat steering working after completed turns
* fix: observe completed progress refresh runs
* test: wait for durable startup recovery admission
Observe committed session rows before stopping startup recovery. Dispatch invocation precedes durable admission, so stopping after a call-count assertion can cancel the write and leave abortedLastRun set. Preserve restored-store capacity sequencing and wait for both stores without assuming recovery order.
Controlled original synchronization reproduces the CI assertion failure; corrected synchronization passes. Ten startup and cancellation cases pass in 26.64 seconds, with owner types, lint, formatting, line-cap guards and independent review through P2 clean.
Use the shared cleanup runner for provider test shards and track xAI fetch mocks so they cannot leak into later transport tests. The same 73-file shard improved from 88.658s to 41.593s across three runs per setting. All four native provider partitions passed: 290 files, 5,234 tests, and one existing opt-in live-test skip.
A test failure without a related change is a defect: re-running, re-pushing,
or refreshing a PR to get green is prohibited, and a failure outside the diff
is not "unrelated" until its cause and owning fix are identified. New or
changed tests state their measured cost and stay within the budgets in the
testing guide, which gains a cost-budget section and a flake-triage recipe.