From 2f4d4751e52bf6fc91fd2a7e77033882a1140ce7 Mon Sep 17 00:00:00 2001 From: Ayaan Zaidi Date: Fri, 2 Oct 2026 01:00:18 +0800 Subject: [PATCH] fix(cron): delegated blocked runs report ok and silent blocked jobs get auto-disabled (#162686) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Related: #162391 ## What Problem This Solves Fixes two follow-ups to #162391: - A scheduled run that hands its work to a subagent can deliver the child's `AUTOMATION_FAILED` reply verbatim and still record the run as `ok`. This happens on announce jobs, and on `delivery: none` jobs since #162474 started recording the child's answer. - A silent (`delivery: none`) job whose agent keeps reporting a blocked task gets auto-disabled after 10 runs, and that posts an auto-disable notice even though the job has nowhere to notify. ## User Impact - When a delegated run's settled child answer starts with `AUTOMATION_FAILED`, the run is recorded as an error with the child's explanation. Announce and current-session delivery send only the explanation, never the token. `delivery: none` runs record it without sending anything, and keep the run transcript as other failed executions do. - A `delivery: none` job with no failure-alert route still records agent-reported failures in run history, job status, and error backoff (escalating to hourly, like any failing job). Those failures never auto-disable the job or post the auto-disable notice, so the job stays silent. - Unchanged: - Runtime errors on silent jobs still auto-disable at 10 as before. - Announce and webhook jobs, and jobs with a configured failure route, still auto-disable on agent-reported failures. - `NO_REPLY` behavior. ## Why This Change Was Made - **Settled child answers.** #162391 classified the token in `resolveCronPayloadOutcome`, which runs on the parent's reply before #117308's descendant settlement. Two places in `dispatchCronDelivery` then adopt the settled child's reply as the run's answer: the announce settlement (#117308) and the no-delivery settlement (#162474). Neither classified it. Both now go through one adoption helper that calls the shared parser `readAutomationFailedReport`, so the run's effective terminal answer is classified after settlement. `buildDeliveryState` derives one execution-failed flag from that result and uses it for both the completion status and the transcript-cleanup decision, so a failed no-delivery run is no longer cleaned up as a quiet success. `run-finalize` consumes the result the same way it already handles `pendingPresentationWarningError`: error status, permanent classification, explanation as the summary. - **Auto-disable.** Maintainer decision: an agent-reported blocked outcome on a job with no notification owner must not lead to auto-disable. The permanent error classification now carries `reportedByAgent`. `applyJobResult` in `timer-outcomes.ts` still counts every error in `consecutiveErrors`, so status and error backoff are unchanged. It skips only the auto-disable decision, and only when all of these hold: the failure was agent-reported, `resolveFailureAlert` finds no route, and the delivery mode is `none` (webhook jobs still auto-disable). - Hermes handles the same protocol the same way (`[CRON_FAILURE]`): its cron hint asks the agent to report a delegated child's failure on the first line. No new tool is needed. ## Bounded cost - No new turns, runs, waits, or notifications. Parsing happens once per settled answer inside the existing finalization. - Agent-reported failures stay permanent, so the scheduler never retries them early. - A silent blocked job keeps the normal error backoff (30 s, 1 min, 5 min, 15 min, then hourly), so it can never run more often than an identical job on origin/main. The only difference is that it keeps running at that backed-off cadence after the 10th failure instead of being disabled. ## Evidence ### Live model, Telegram Test Server (`telegram-e2e-userbot`, DM) Setup: the repo runner unchanged (`run-mock-sut-user-e2e.mjs --backend mock --source-gateway --dm`, Convex lease). `E2E_MOCK_SERVER_PATH` points to a throwaway proxy that stands in for the mock provider and forwards every request unchanged to the real OpenAI API. The proxy reads the key from a private file, so the Gateway only ever holds the runner's dummy key. `E2E_ROOT_CONFIG_PATCH` selects `openai/gpt-6-astra` (`tools.toolSearch: false`). A scenario `command` step creates the jobs with `openclaw automations add` and runs them. No model output was steered: the live parent chose `sessions_spawn`, and the live child and the live silent job wrote their own `AUTOMATION_FAILED` reports because they had no shell tool. The only steer is timing in case (b): before each run, `cron.update` sets `nextRunAtMs` to now and moves `lastRunAtMs` back 2 h, so the error backoff (up to 1 h) counts as served. The Gateway's own scheduler then runs the job as a normal scheduled (not forced) run. Before = origin/main `d8a110b1a15`, plus this PR's exact base `5c4e9aba929` for case (a). After = this head `8b77715ba4f`, with all three cases in one Gateway run. **(a) Announce job, live parent delegates to one child, the child reports `AUTOMATION_FAILED`** - before (`d8a110b1a15`): `cron run --wait` took 99.4 s. Run `ok`, `completionStatus: succeeded`, `deliveryStatus: delivered`. The DM received bot message 43601: `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.` The token is visible. - before (`5c4e9aba929`, the PR base): 80.7 s. Same result: run `ok`, `delivered`, and bot message 43705 reads `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.` - after: 63.5 s. Run `error`, `completionStatus: failed`, `deliveryStatus: delivered`. Error and summary are both the explanation. The DM received bot message 43712 with only `The command could not run because no shell execution tool is available in this session.` No token. - Pre-existing flake, not in this diff: on both sides, the announce settlement sometimes ends with `cron child-session handoff completed without a final assistant payload` even though the live child answered. That happened on 1 of 3 origin/main-side attempts (a second `d8a110b1a15` run) and on 2 of 4 attempts on this branch (both on `bfe1f0f2169`). In those cases nothing reaches the chat. The PR's change only runs after a reply has been read, so it is not involved. The no-delivery path (case c) settled on every attempt. **(c) Same delegated job with `delivery: none` (the #162474 path)** - before: 55.2 s. Run `ok`, `succeeded`, `deliveryStatus: not-requested`, summary `AUTOMATION_FAILED ⏎ The command could not run because no shell execution tool is available in this session.` The token is recorded as an ok answer. - after: 36.3 s. Run `error`, `failed`, `not-requested`. Error and summary are both `The command could not run because no shell execution tool is available in this session.` The job stays enabled. No chat message. **(b) `delivery: none` job, no failure-alert route, the live model reports `AUTOMATION_FAILED` on every run (11 scheduled runs)** - before: runs 1–10 were all `error` (`Could not run \`node scripts/ledger-sync.mjs\`: this session only provides a file-reading tool…`), with `consecutiveErrors` 1…10. After run 10 the job was `enabled: false` with `autoDisabled: {reason: "consecutive-failures", consecutiveErrors: 10}`, so run 11 never happened. The DM received bot message 43644 (heartbeat relay): `⚠️ Your "Nightly ledger sync" automation was automatically disabled after 10 consecutive failures. … openclaw automations enable `. 10 runs took 504 s. - after: all 11 runs were `error` with the same explanation, and `consecutiveErrors` went 1…11, which keeps feeding backoff. After run 11 the job was still `enabled: true` with no `autoDisabled`. History shows 11 `error` entries. Across the whole run, the DM received exactly one bot message: 43712 from case (a). Cases (c) and (b) sent none, including during a 120 s wait after the loop. 11 runs took 526 s. Provider evidence (proxy log): before had 1 parent `sessions_spawn` turn, 1 child turn, 10 silent-job turns, and 1 relay turn carrying the auto-disable notice. After had 2 parent `sessions_spawn` turns, 2 child turns, 11 silent-job turns, no auto-disable relay turn, and only one routine heartbeat poll, which answered `NO_REPLY`. Environment note: in `--source-gateway` mode on both refs, Telegram inbound polling could not start (the ingress worker's plugin capture has no `dist/plugin-sdk`). So the opening DM turn was not processed, and the jobs are owned by `agent:main:main`, the DM's owner session under the default `dmScope`. Outbound delivery and relays reached the DM normally. This is the same for before and after. ### Tests Each new test fails on the code it fixes and passes on this head: - `delivery-dispatch.double-announce.test.ts`, real `dispatchCronDelivery`: - Announce settlement: a settled child reply `AUTOMATION_FAILED ⏎ …` is delivered as the explanation only and recorded as `agentReportedFailure`. - No-delivery settlement: uses the real `waitForDescendantSubagentResult`, with the registry mocked only at its edge. The child's report is recorded as `agentReportedFailure` with the explanation as output and summary. Nothing is sent, and the `deleteAfterRun` transcript is not deleted. Classification fails without the shared adoption helper (checked on the rebased predecessor commit), and retention fails on `bfe1f0f2169`. - `timer-outcomes.regression.test.ts`, real `applyJobResult`, 11 agent-reported failures from a clean streak: - `delivery: none`: still enabled, streak 11, no auto-disable notice, backoff delays 5 min, 15 min, 1 h, 1 h. This fails on `ee6d9e07fdf` (no streak, no backoff). - Webhook and announce jobs: auto-disabled at 10 with one notice, same backoff. Measured on this head, run separately (vitest `Duration`): - `delivery-dispatch.double-announce`: 84 tests, 22.1 s. - `timer-outcomes.regression`: 20 tests, 15.5 s. - `run.meta-error-status`: 14 tests, 20.0 s. All pass. `pnpm tsgo:core` passes, and oxlint and oxfmt on the changed files are clean. Docs: `delivery.md` documents the silent agent-failure exception to the 10-failure auto-disable. `payloads.md` notes that a delegated run's settled answer is classified the same way. LOC vs origin/main: production +54/−25, tests +100/−1, docs +2/−2. ### Review - Adversarial re-review round 1 (on `0530dc9f`): fixed the one real finding. Webhook jobs had been exempted too; the exemption now requires delivery mode `none`, with a webhook test case. - Round 2 (on `ee6d9e07fdf`): NOT LAND, one finding. Skipping the increment also froze error backoff, so a silent blocked job could keep running every minute. Fixed per maintainer decision with no second counter: the streak always increments, and the exemption moved to the auto-disable decision. - Round 3 (on `bfe1f0f2169`, rebased): LAND, no findings. The rebase onto origin/main also routes the #162474 no-delivery settlement through the same adoption helper. - ClawSweeper on `bfe1f0f2169`: fixed both findings. (P2) A failed no-delivery child run deleted its transcript; it is now kept. (P3) The auto-disable exception is now documented. Also added per-file test timings. - Round 4 (on `8b77715ba4f`): LAND, no findings. - Accepted tradeoffs: the exemption is not re-derived during finalized-run startup recovery. A silent job whose streak is already 10 or more from reported failures auto-disables on its next runtime error. Co-authored-by: Ayaan Zaidi --- docs/automation/cron-jobs/delivery.md | 2 +- docs/automation/cron-jobs/payloads.md | 2 +- .../isolated-agent/delivery-dispatch-types.ts | 2 + .../delivery-dispatch.double-announce.test.ts | 37 +++++++++++ src/cron/isolated-agent/delivery-dispatch.ts | 29 +++++---- src/cron/isolated-agent/helpers.ts | 25 ++++---- src/cron/isolated-agent/run-finalize.ts | 10 ++- .../run.meta-error-status.test.ts | 2 +- .../service/timer-outcomes.regression.test.ts | 62 +++++++++++++++++++ src/cron/service/timer-outcomes.ts | 10 +++ src/cron/types.ts | 3 +- 11 files changed, 156 insertions(+), 28 deletions(-) diff --git a/docs/automation/cron-jobs/delivery.md b/docs/automation/cron-jobs/delivery.md index d6f8a1674dbe..e073e560731c 100644 --- a/docs/automation/cron-jobs/delivery.md +++ b/docs/automation/cron-jobs/delivery.md @@ -148,7 +148,7 @@ Script setup refreshes retired tools after a plugin reload before execution begi A provider rejection of an unsupported model records `model_not_found` in the job state and run history. The failure notice points to `openclaw doctor --fix` for provider-declared retirements, or changing/removing the automation's model override. Known retired automation model routes fail before another inference request. Doctor replaces an override with the provider's declared successor when the agent's model policy allows it. Without a declared successor, it clears the override so the job inherits the agent default. If a pinned override's successor is disallowed, Doctor retains the reference and reports the required policy change. A missing account catalog entry or a discovery outage alone does not authorize a migration. -The scheduler also provides an unconditional safety backstop. A time-based recurring job is auto-disabled after 10 consecutive execution failures; a successful run resets that streak. On the terminal failure, the richer auto-disable notification replaces the regular threshold alert. Repeated schedule-computation failures auto-disable after 3 errors. The job records `state.autoDisabled.reason` as `consecutive-failures` or `schedule-errors`, and the owning agent receives a notification with a safe cause and recovery command. Raw errors stay in automation history. After fixing the cause, run `openclaw automations enable `; enabling clears the recorded reason and failure streaks. Because disabled jobs are hidden by the default list, use `openclaw automations list --all` to inspect them. +The scheduler also provides a safety backstop. A time-based recurring job is auto-disabled after 10 consecutive execution failures; a successful run resets that streak. One exception: a `delivery.mode: "none"` job with no active failure-alert policy is never auto-disabled by its agent's own `AUTOMATION_FAILED` reports, because it has no one to notify. Those runs still count toward `consecutiveErrors` and the error backoff, and its runtime errors still auto-disable it. On the terminal failure, the richer auto-disable notification replaces the regular threshold alert. Repeated schedule-computation failures auto-disable after 3 errors. The job records `state.autoDisabled.reason` as `consecutive-failures` or `schedule-errors`, and the owning agent receives a notification with a safe cause and recovery command. Raw errors stay in automation history. After fixing the cause, run `openclaw automations enable `; enabling clears the recorded reason and failure streaks. Because disabled jobs are hidden by the default list, use `openclaw automations list --all` to inspect them. ### Output language diff --git a/docs/automation/cron-jobs/payloads.md b/docs/automation/cron-jobs/payloads.md index 73fa5b567276..de1477ad28d7 100644 --- a/docs/automation/cron-jobs/payloads.md +++ b/docs/automation/cron-jobs/payloads.md @@ -303,7 +303,7 @@ Agent-turn jobs default to the creating conversation when the create request car A new transcript/session id per run. OpenClaw carries safe preferences (thinking/fast/verbose settings, labels, explicit user-selected model/auth overrides), but does not inherit ambient conversation context from an older automation session row: channel/group routing, send or queue policy, elevation, origin, or ACP runtime binding. Use `current` or `session:` when a recurring job should deliberately build on the same conversation context. - Isolated automation and hook agent turns are explicitly unattended: no one is present to clarify or approve. The final reply must be the deliverable rather than a plan, acknowledgement, or request for input. The agent returns `NO_REPLY` when nothing needs doing. When the task failed or is blocked, the reply starts with `AUTOMATION_FAILED` on its own line, followed by what failed and what it tried. The scheduler records that run as an error with the remaining text as its error, delivers that text instead of the token when the job announces, and applies the normal retry, failure-alert, and owner-repair policy. Only an exact first line counts; a reply that mentions the token elsewhere is ordinary output. + Isolated automation and hook agent turns are explicitly unattended: no one is present to clarify or approve. The final reply must be the deliverable rather than a plan, acknowledgement, or request for input. The agent returns `NO_REPLY` when nothing needs doing. When the task failed or is blocked, the reply starts with `AUTOMATION_FAILED` on its own line, followed by what failed and what it tried. The scheduler records that run as an error with the remaining text as its error, delivers that text instead of the token when the job announces, and applies the normal retry, failure-alert, and owner-repair policy. When the run hands its work to a subagent, the child's settled final answer is classified the same way. Only an exact first line counts; a reply that mentions the token elsewhere is ordinary output. For trusted scheduled jobs, the job's own instructions win when they intentionally ask for a question or plan, and the agent may remove a job that is no longer needed. External hook turns receive only the common unattended contract; they do not receive that override or self-removal guidance across the external-content boundary. diff --git a/src/cron/isolated-agent/delivery-dispatch-types.ts b/src/cron/isolated-agent/delivery-dispatch-types.ts index 0585c808eabb..44680359ef99 100644 --- a/src/cron/isolated-agent/delivery-dispatch-types.ts +++ b/src/cron/isolated-agent/delivery-dispatch-types.ts @@ -66,4 +66,6 @@ export type DispatchCronDeliveryState = { outputText?: string; synthesizedText?: string; deliveryPayloads: ReplyPayload[]; + /** Explanation from a settled descendant answer that reported AUTOMATION_FAILED. */ + agentReportedFailure?: string; }; diff --git a/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts b/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts index 720ec521b277..3401360fdd3c 100644 --- a/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts +++ b/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts @@ -733,6 +733,22 @@ describe("dispatchCronDelivery", () => { expect(state.deliveryError).toBe("cron descendants completed without a final reply"); }); + it("classifies a settled child's AUTOMATION_FAILED answer and delivers only its explanation", async () => { + vi.mocked(readDescendantSubagentFallbackReply).mockResolvedValue( + "AUTOMATION_FAILED\nNo shell tool is available in this run.", + ); + + const state = await dispatchCronDelivery(emptyParams(true)); + + expect(deliverOutboundPayloads).toHaveBeenCalledTimes(1); + expectDeliveryCall(0, { payloads: [{ text: "No shell tool is available in this run." }] }); + expect(state).toMatchObject({ + delivered: true, + agentReportedFailure: "No shell tool is available in this run.", + summary: "No shell tool is available in this run.", + }); + }); + it.each([ ["active threaded best-effort", true, "42", true], ["completed direct", false, undefined, false], @@ -891,6 +907,27 @@ describe("dispatchCronDelivery", () => { expectSessionDeleted(); }); + it("records a child's AUTOMATION_FAILED report as a failure and keeps the transcript", async () => { + const params = spawnOnlyJob({ mode: "none" }); + childSettlesAt(1_000, { + disposition: "visible", + text: "AUTOMATION_FAILED\nNo shell tool is available in this run.", + }); + + const state = await dispatchUntilWatchdog(params); + + expect(state).toMatchObject({ + agentReportedFailure: "No shell tool is available in this run.", + outputText: "No shell tool is available in this run.", + summary: "No shell tool is available in this run.", + deliveryState: { status: "not-requested" }, + }); + expect(deliverOutboundPayloads).not.toHaveBeenCalled(); + expect(callGateway).not.toHaveBeenCalledWith( + expect.objectContaining({ method: "sessions.delete" }), + ); + }); + it.each([ { name: "the child never settles", diff --git a/src/cron/isolated-agent/delivery-dispatch.ts b/src/cron/isolated-agent/delivery-dispatch.ts index 491a258b76b1..71edc6bce5d0 100644 --- a/src/cron/isolated-agent/delivery-dispatch.ts +++ b/src/cron/isolated-agent/delivery-dispatch.ts @@ -61,7 +61,7 @@ import { appendCronRunInspectionLink, normalizeDirectCronDeliveryPayloads, } from "./delivery-payload-normalization.js"; -import { pickSummaryFromOutput } from "./helpers.js"; +import { pickSummaryFromOutput, readAutomationFailedReport } from "./helpers.js"; import { cleanupCronRunSessionAfterRun } from "./session-cleanup.js"; import { isLikelyInterimCronMessage } from "./subagent-followup-hints.js"; @@ -83,6 +83,17 @@ export async function dispatchCronDelivery( let outputText = params.outputText; let synthesizedText = params.synthesizedText; let deliveryPayloads = params.deliveryPayloads; + let agentReportedFailure: string | undefined; + // A settled descendant answer is the run's terminal answer; classify it like the parent's + // own reply so a reported failure is recorded as one and never delivered as a token. + const adoptSettledChildReply = (childReply: string) => { + agentReportedFailure = readAutomationFailedReport(childReply); + const reply = agentReportedFailure ?? childReply; + outputText = reply; + summary = pickSummaryFromOutput(reply) ?? summary; + synthesizedText = reply; + deliveryPayloads = [{ text: reply }]; + }; const deliveryState: CronResolvedDeliveryState = { status: params.deliveryRequested ? "not-delivered" : "not-requested", @@ -106,16 +117,17 @@ export async function dispatchCronDelivery( let deliveryAttempted = verifiedMessageToolDelivery; let deferredDeletingSessionMirror: DirectCronTranscriptMirror | undefined; const buildDeliveryState = async (disposition?: CronDeliveryDisposition) => { + const executionFailed = disposition?.kind === "error" || agentReportedFailure !== undefined; const completion = resolveAdmittedCronCompletionStatus( params.job, - disposition?.kind === "error" ? "error" : params.undeliveredRunStatus, + executionFailed ? "error" : params.undeliveredRunStatus, deliveryState.status, deliveryState.deliverySuppressionReason, ); // Quiet/best-effort successes retire with their jobs; failed executions retain evidence. if ( deliveryState.status === "delivered" || - (deliveryState.status === "not-requested" && disposition?.kind !== "error") || + (deliveryState.status === "not-requested" && !executionFailed) || completion === "succeeded" ) { await cleanupDirectCronSessionIfNeeded(); @@ -132,6 +144,7 @@ export async function dispatchCronDelivery( outputText, synthesizedText, deliveryPayloads, + ...(agentReportedFailure ? { agentReportedFailure } : {}), }; }; const formatDeliveryTargetError = (error: string) => @@ -556,10 +569,7 @@ export async function dispatchCronDelivery( abortSignal: params.abortSignal, }); if (finalReply) { - outputText = finalReply; - summary = pickSummaryFromOutput(finalReply) ?? summary; - synthesizedText = finalReply; - deliveryPayloads = [{ text: finalReply }]; + adoptSettledChildReply(finalReply); } if (spawnOnlyHandoff && !synthesizedText?.trim()) { // An accepted spawn is the turn's only completion; retiring it without @@ -724,10 +734,7 @@ export async function dispatchCronDelivery( }); } if (!isSilentReplyText(settled.reply, SILENT_REPLY_TOKEN)) { - outputText = settled.reply; - summary = pickSummaryFromOutput(settled.reply) ?? summary; - synthesizedText = settled.reply; - deliveryPayloads = [{ text: settled.reply }]; + adoptSettledChildReply(settled.reply); } } diff --git a/src/cron/isolated-agent/helpers.ts b/src/cron/isolated-agent/helpers.ts index 885d6f08e903..91801ce252a8 100644 --- a/src/cron/isolated-agent/helpers.ts +++ b/src/cron/isolated-agent/helpers.ts @@ -70,6 +70,17 @@ function formatCronRunLevelError(error: unknown): string | undefined { return detail ? `cron isolated run failed: ${detail}` : "cron isolated run failed"; } +/** + * Returns the explanation when a reply reports AUTOMATION_FAILED on its exact first line. + * A reply that only quotes the token elsewhere stays ordinary output. + */ +export function readAutomationFailedReport(text: string | undefined): string | undefined { + const [firstLine, ...detail] = (text ?? "").trim().split("\n"); + return firstLine?.trim() === AUTOMATION_FAILED_TOKEN + ? (normalizeOptionalString(detail.join("\n")) ?? "The automation run reported that it failed.") + : undefined; +} + /** Picks a bounded cron run summary from plain text output. */ export function pickSummaryFromOutput(text: string | undefined) { const clean = (text ?? "").trim(); @@ -293,17 +304,9 @@ export function resolveCronPayloadOutcome(params: { ? { ...params.failureSignal, message: failureMessage } : undefined; const runLevelError = formatCronRunLevelError(params.runLevelError); - // Only an exact first line counts; a reply that quotes the token stays ordinary output. - const [reportedFirstLine, ...reportedDetail] = ( - normalizedFinalAssistantVisibleText ?? - fallbackOutputText ?? - "" - ).split("\n"); - const reportedFailure = - reportedFirstLine?.trim() === AUTOMATION_FAILED_TOKEN - ? (normalizeOptionalString(reportedDetail.join("\n")) ?? - "The automation run reported that it failed.") - : undefined; + const reportedFailure = readAutomationFailedReport( + normalizedFinalAssistantVisibleText ?? fallbackOutputText, + ); const hasFatalErrorPayload = hasFatalStructuredErrorPayload || failureSignal !== undefined || diff --git a/src/cron/isolated-agent/run-finalize.ts b/src/cron/isolated-agent/run-finalize.ts index 4f590c828baf..7678cfb517cb 100644 --- a/src/cron/isolated-agent/run-finalize.ts +++ b/src/cron/isolated-agent/run-finalize.ts @@ -227,6 +227,7 @@ export async function finalizeCronRun(params: { outputText, hasFatalErrorPayload, embeddedRunError, + agentReportedFailure, } = cronPayloadOutcome; const terminalToolFailure = finalRunResult.meta?.terminalToolFailure; const hasTerminalToolFailure = isEmbeddedRunTerminalToolFailure(terminalToolFailure); @@ -262,8 +263,8 @@ export async function finalizeCronRun(params: { error: runError, // The agent already judged the task blocked: rerunning it would repeat that turn, // and its prose must not be text-classified into a transient retry reason. - ...(cronPayloadOutcome.agentReportedFailure - ? { errorClassification: { kind: "permanent" as const } } + ...(agentReportedFailure + ? { errorClassification: { kind: "permanent" as const, reportedByAgent: true } } : {}), } : {}), @@ -412,5 +413,10 @@ export async function finalizeCronRun(params: { hasFatalErrorPayload = true; embeddedRunError = pendingPresentationWarningError; } + if (deliveryResult.agentReportedFailure && !hasFatalErrorPayload) { + hasFatalErrorPayload = true; + embeddedRunError = deliveryResult.agentReportedFailure; + agentReportedFailure = true; + } return resolveRunOutcome({ ...deliveryResult, delivery: deliveryTrace }); } diff --git a/src/cron/isolated-agent/run.meta-error-status.test.ts b/src/cron/isolated-agent/run.meta-error-status.test.ts index c726e2291409..52e19c3980f4 100644 --- a/src/cron/isolated-agent/run.meta-error-status.test.ts +++ b/src/cron/isolated-agent/run.meta-error-status.test.ts @@ -168,7 +168,7 @@ describe("runCronIsolatedAgentTurn - meta.error status propagation", () => { expected: { status: "error", error: "Network timeout: no shell tool is available in this run.", - errorClassification: { kind: "permanent" }, + errorClassification: { kind: "permanent", reportedByAgent: true }, }, }, { diff --git a/src/cron/service/timer-outcomes.regression.test.ts b/src/cron/service/timer-outcomes.regression.test.ts index c70a3ebbfd28..b555f56b8faf 100644 --- a/src/cron/service/timer-outcomes.regression.test.ts +++ b/src/cron/service/timer-outcomes.regression.test.ts @@ -174,6 +174,68 @@ describe("cron timer outcome and failure policy regressions", () => { expect(sendCronFailureAlert).not.toHaveBeenCalled(); }); + it.each([ + { name: "silent job", delivery: { mode: "none" as const }, disables: false }, + { + name: "webhook job", + delivery: { mode: "webhook" as const, to: "https://hooks.example.test/cron" }, + disables: true, + }, + { name: "announce job", delivery: undefined, disables: true }, + ])( + "backs off agent-reported failures and auto-disables only with a notification owner: $name", + ({ delivery, disables }) => { + const startedAt = Date.parse("2026-08-01T12:00:00.000Z"); + let now = startedAt; + const deferredNotifications: DeferredCronNotifications = []; + const state = createCronServiceState({ + storePath: "/tmp/cron-reported-failure-threshold.json", + nowMs: () => now, + enqueueSystemEvent: vi.fn(), + runIsolatedAgentJob: createDefaultIsolatedRunner(), + }); + const job = createIsolatedRegressionJob({ + id: "recurring-reported-failure", + name: "recurring reported failure", + scheduledAt: startedAt, + schedule: { kind: "every", everyMs: 60_000, anchorMs: startedAt }, + payload: { kind: "agentTurn", message: "report" }, + state: {}, + }); + job.delivery = delivery; + + const delaysMs: number[] = []; + for (let run = 1; run <= 11 && job.enabled; run += 1) { + applyJobResult( + state, + job, + { + status: "error", + error: "No shell tool is available in this run.", + errorClassification: { kind: "permanent", reportedByAgent: true }, + startedAt: now, + endedAt: now + 10, + }, + { deferredNotifications }, + ); + if (job.state.nextRunAtMs !== undefined) { + delaysMs.push(job.state.nextRunAtMs - (now + 10)); + now = job.state.nextRunAtMs; + } + } + + expect(job.state.lastRunStatus).toBe("error"); + expect(job.state.lastError).toBe("No shell tool is available in this run."); + expect(job.enabled).toBe(!disables); + expect(job.state.consecutiveErrors).toBe(disables ? 10 : 11); + expect(deferredNotifications.filter((n) => n.kind === "auto-disabled")).toHaveLength( + disables ? 1 : 0, + ); + // The streak still drives error backoff past the one-minute cadence: 5 and 15 minutes, then hourly. + expect(delaysMs.slice(2, 6)).toEqual([300_000, 900_000, 3_600_000, 3_600_000]); + }, + ); + it("resets the auto-disable streak after a successful recurring run", () => { const startedAt = Date.parse("2026-08-01T13:00:00.000Z"); const state = createCronServiceState({ diff --git a/src/cron/service/timer-outcomes.ts b/src/cron/service/timer-outcomes.ts index 451c09d5aaa3..34c7dcaed89a 100644 --- a/src/cron/service/timer-outcomes.ts +++ b/src/cron/service/timer-outcomes.ts @@ -1,6 +1,7 @@ import { resolveCronTriggerMinIntervalMs } from "../../config/cron-limits.js"; import type { CronActiveJobMarker } from "../active-jobs.js"; import { resolveAdmittedCronCompletionStatus } from "../completion-status.js"; +import { resolveCronDeliveryPlan } from "../delivery-plan.js"; import { resolvePacedNextRunAtMs } from "../pacing.js"; import { normalizeCronRunDiagnostics, summarizeCronRunDiagnostics } from "../run-diagnostics.js"; import { resolveCronRunErrorReason } from "../run-error-reason.js"; @@ -185,6 +186,14 @@ export function applyJobResult( } }; const alertConfig = resolveFailureAlert(state, job); + // A silent job's agent-reported blocked outcome stays in history, status, and backoff, but no + // notification owner exists for it, so it never auto-disables the job or posts that notice. + const silentReportedFailure = + result.status === "error" && + result.errorClassification?.kind === "permanent" && + result.errorClassification.reportedByAgent === true && + alertConfig === null && + resolveCronDeliveryPlan(job).mode === "none"; if (result.status === "error") { job.state.consecutiveErrors = (job.state.consecutiveErrors ?? 0) + 1; job.state.consecutiveSkipped = 0; @@ -352,6 +361,7 @@ export function applyJobResult( } else if ( result.status === "error" && isJobEnabled(job) && + !silentReportedFailure && maybeAutoDisableCronJobAfterRunFailure({ job, atMs: result.endedAt, diff --git a/src/cron/types.ts b/src/cron/types.ts index 0acde0532b79..fcc897bde904 100644 --- a/src/cron/types.ts +++ b/src/cron/types.ts @@ -151,7 +151,8 @@ export type CronRunDiagnostics = NonNullable /** Explicit execution-error disposition used consistently by retry, history, and alerts. */ export type CronRunErrorClassification = | { kind: "reason"; reason: FailoverReason } - | { kind: "permanent" }; + /** `reportedByAgent`: the run's final answer reported AUTOMATION_FAILED; no runtime fault. */ + | { kind: "permanent"; reportedByAgent?: true }; /** Closed producer-authored facts allowed in operator-facing failure notifications. */ export type CronFailureNotificationDetail =