diff --git a/docs/automation/cron-jobs/delivery.md b/docs/automation/cron-jobs/delivery.md index d6f8a1674dbe..e073e560731c 100644 --- a/docs/automation/cron-jobs/delivery.md +++ b/docs/automation/cron-jobs/delivery.md @@ -148,7 +148,7 @@ Script setup refreshes retired tools after a plugin reload before execution begi A provider rejection of an unsupported model records `model_not_found` in the job state and run history. The failure notice points to `openclaw doctor --fix` for provider-declared retirements, or changing/removing the automation's model override. Known retired automation model routes fail before another inference request. Doctor replaces an override with the provider's declared successor when the agent's model policy allows it. Without a declared successor, it clears the override so the job inherits the agent default. If a pinned override's successor is disallowed, Doctor retains the reference and reports the required policy change. A missing account catalog entry or a discovery outage alone does not authorize a migration. -The scheduler also provides an unconditional safety backstop. A time-based recurring job is auto-disabled after 10 consecutive execution failures; a successful run resets that streak. On the terminal failure, the richer auto-disable notification replaces the regular threshold alert. Repeated schedule-computation failures auto-disable after 3 errors. The job records `state.autoDisabled.reason` as `consecutive-failures` or `schedule-errors`, and the owning agent receives a notification with a safe cause and recovery command. Raw errors stay in automation history. After fixing the cause, run `openclaw automations enable `; enabling clears the recorded reason and failure streaks. Because disabled jobs are hidden by the default list, use `openclaw automations list --all` to inspect them. +The scheduler also provides a safety backstop. A time-based recurring job is auto-disabled after 10 consecutive execution failures; a successful run resets that streak. One exception: a `delivery.mode: "none"` job with no active failure-alert policy is never auto-disabled by its agent's own `AUTOMATION_FAILED` reports, because it has no one to notify. Those runs still count toward `consecutiveErrors` and the error backoff, and its runtime errors still auto-disable it. On the terminal failure, the richer auto-disable notification replaces the regular threshold alert. Repeated schedule-computation failures auto-disable after 3 errors. The job records `state.autoDisabled.reason` as `consecutive-failures` or `schedule-errors`, and the owning agent receives a notification with a safe cause and recovery command. Raw errors stay in automation history. After fixing the cause, run `openclaw automations enable `; enabling clears the recorded reason and failure streaks. Because disabled jobs are hidden by the default list, use `openclaw automations list --all` to inspect them. ### Output language diff --git a/docs/automation/cron-jobs/payloads.md b/docs/automation/cron-jobs/payloads.md index 73fa5b567276..de1477ad28d7 100644 --- a/docs/automation/cron-jobs/payloads.md +++ b/docs/automation/cron-jobs/payloads.md @@ -303,7 +303,7 @@ Agent-turn jobs default to the creating conversation when the create request car A new transcript/session id per run. OpenClaw carries safe preferences (thinking/fast/verbose settings, labels, explicit user-selected model/auth overrides), but does not inherit ambient conversation context from an older automation session row: channel/group routing, send or queue policy, elevation, origin, or ACP runtime binding. Use `current` or `session:` when a recurring job should deliberately build on the same conversation context. - Isolated automation and hook agent turns are explicitly unattended: no one is present to clarify or approve. The final reply must be the deliverable rather than a plan, acknowledgement, or request for input. The agent returns `NO_REPLY` when nothing needs doing. When the task failed or is blocked, the reply starts with `AUTOMATION_FAILED` on its own line, followed by what failed and what it tried. The scheduler records that run as an error with the remaining text as its error, delivers that text instead of the token when the job announces, and applies the normal retry, failure-alert, and owner-repair policy. Only an exact first line counts; a reply that mentions the token elsewhere is ordinary output. + Isolated automation and hook agent turns are explicitly unattended: no one is present to clarify or approve. The final reply must be the deliverable rather than a plan, acknowledgement, or request for input. The agent returns `NO_REPLY` when nothing needs doing. When the task failed or is blocked, the reply starts with `AUTOMATION_FAILED` on its own line, followed by what failed and what it tried. The scheduler records that run as an error with the remaining text as its error, delivers that text instead of the token when the job announces, and applies the normal retry, failure-alert, and owner-repair policy. When the run hands its work to a subagent, the child's settled final answer is classified the same way. Only an exact first line counts; a reply that mentions the token elsewhere is ordinary output. For trusted scheduled jobs, the job's own instructions win when they intentionally ask for a question or plan, and the agent may remove a job that is no longer needed. External hook turns receive only the common unattended contract; they do not receive that override or self-removal guidance across the external-content boundary. diff --git a/src/cron/isolated-agent/delivery-dispatch-types.ts b/src/cron/isolated-agent/delivery-dispatch-types.ts index 0585c808eabb..44680359ef99 100644 --- a/src/cron/isolated-agent/delivery-dispatch-types.ts +++ b/src/cron/isolated-agent/delivery-dispatch-types.ts @@ -66,4 +66,6 @@ export type DispatchCronDeliveryState = { outputText?: string; synthesizedText?: string; deliveryPayloads: ReplyPayload[]; + /** Explanation from a settled descendant answer that reported AUTOMATION_FAILED. */ + agentReportedFailure?: string; }; diff --git a/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts b/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts index 720ec521b277..3401360fdd3c 100644 --- a/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts +++ b/src/cron/isolated-agent/delivery-dispatch.double-announce.test.ts @@ -733,6 +733,22 @@ describe("dispatchCronDelivery", () => { expect(state.deliveryError).toBe("cron descendants completed without a final reply"); }); + it("classifies a settled child's AUTOMATION_FAILED answer and delivers only its explanation", async () => { + vi.mocked(readDescendantSubagentFallbackReply).mockResolvedValue( + "AUTOMATION_FAILED\nNo shell tool is available in this run.", + ); + + const state = await dispatchCronDelivery(emptyParams(true)); + + expect(deliverOutboundPayloads).toHaveBeenCalledTimes(1); + expectDeliveryCall(0, { payloads: [{ text: "No shell tool is available in this run." }] }); + expect(state).toMatchObject({ + delivered: true, + agentReportedFailure: "No shell tool is available in this run.", + summary: "No shell tool is available in this run.", + }); + }); + it.each([ ["active threaded best-effort", true, "42", true], ["completed direct", false, undefined, false], @@ -891,6 +907,27 @@ describe("dispatchCronDelivery", () => { expectSessionDeleted(); }); + it("records a child's AUTOMATION_FAILED report as a failure and keeps the transcript", async () => { + const params = spawnOnlyJob({ mode: "none" }); + childSettlesAt(1_000, { + disposition: "visible", + text: "AUTOMATION_FAILED\nNo shell tool is available in this run.", + }); + + const state = await dispatchUntilWatchdog(params); + + expect(state).toMatchObject({ + agentReportedFailure: "No shell tool is available in this run.", + outputText: "No shell tool is available in this run.", + summary: "No shell tool is available in this run.", + deliveryState: { status: "not-requested" }, + }); + expect(deliverOutboundPayloads).not.toHaveBeenCalled(); + expect(callGateway).not.toHaveBeenCalledWith( + expect.objectContaining({ method: "sessions.delete" }), + ); + }); + it.each([ { name: "the child never settles", diff --git a/src/cron/isolated-agent/delivery-dispatch.ts b/src/cron/isolated-agent/delivery-dispatch.ts index 491a258b76b1..71edc6bce5d0 100644 --- a/src/cron/isolated-agent/delivery-dispatch.ts +++ b/src/cron/isolated-agent/delivery-dispatch.ts @@ -61,7 +61,7 @@ import { appendCronRunInspectionLink, normalizeDirectCronDeliveryPayloads, } from "./delivery-payload-normalization.js"; -import { pickSummaryFromOutput } from "./helpers.js"; +import { pickSummaryFromOutput, readAutomationFailedReport } from "./helpers.js"; import { cleanupCronRunSessionAfterRun } from "./session-cleanup.js"; import { isLikelyInterimCronMessage } from "./subagent-followup-hints.js"; @@ -83,6 +83,17 @@ export async function dispatchCronDelivery( let outputText = params.outputText; let synthesizedText = params.synthesizedText; let deliveryPayloads = params.deliveryPayloads; + let agentReportedFailure: string | undefined; + // A settled descendant answer is the run's terminal answer; classify it like the parent's + // own reply so a reported failure is recorded as one and never delivered as a token. + const adoptSettledChildReply = (childReply: string) => { + agentReportedFailure = readAutomationFailedReport(childReply); + const reply = agentReportedFailure ?? childReply; + outputText = reply; + summary = pickSummaryFromOutput(reply) ?? summary; + synthesizedText = reply; + deliveryPayloads = [{ text: reply }]; + }; const deliveryState: CronResolvedDeliveryState = { status: params.deliveryRequested ? "not-delivered" : "not-requested", @@ -106,16 +117,17 @@ export async function dispatchCronDelivery( let deliveryAttempted = verifiedMessageToolDelivery; let deferredDeletingSessionMirror: DirectCronTranscriptMirror | undefined; const buildDeliveryState = async (disposition?: CronDeliveryDisposition) => { + const executionFailed = disposition?.kind === "error" || agentReportedFailure !== undefined; const completion = resolveAdmittedCronCompletionStatus( params.job, - disposition?.kind === "error" ? "error" : params.undeliveredRunStatus, + executionFailed ? "error" : params.undeliveredRunStatus, deliveryState.status, deliveryState.deliverySuppressionReason, ); // Quiet/best-effort successes retire with their jobs; failed executions retain evidence. if ( deliveryState.status === "delivered" || - (deliveryState.status === "not-requested" && disposition?.kind !== "error") || + (deliveryState.status === "not-requested" && !executionFailed) || completion === "succeeded" ) { await cleanupDirectCronSessionIfNeeded(); @@ -132,6 +144,7 @@ export async function dispatchCronDelivery( outputText, synthesizedText, deliveryPayloads, + ...(agentReportedFailure ? { agentReportedFailure } : {}), }; }; const formatDeliveryTargetError = (error: string) => @@ -556,10 +569,7 @@ export async function dispatchCronDelivery( abortSignal: params.abortSignal, }); if (finalReply) { - outputText = finalReply; - summary = pickSummaryFromOutput(finalReply) ?? summary; - synthesizedText = finalReply; - deliveryPayloads = [{ text: finalReply }]; + adoptSettledChildReply(finalReply); } if (spawnOnlyHandoff && !synthesizedText?.trim()) { // An accepted spawn is the turn's only completion; retiring it without @@ -724,10 +734,7 @@ export async function dispatchCronDelivery( }); } if (!isSilentReplyText(settled.reply, SILENT_REPLY_TOKEN)) { - outputText = settled.reply; - summary = pickSummaryFromOutput(settled.reply) ?? summary; - synthesizedText = settled.reply; - deliveryPayloads = [{ text: settled.reply }]; + adoptSettledChildReply(settled.reply); } } diff --git a/src/cron/isolated-agent/helpers.ts b/src/cron/isolated-agent/helpers.ts index 885d6f08e903..91801ce252a8 100644 --- a/src/cron/isolated-agent/helpers.ts +++ b/src/cron/isolated-agent/helpers.ts @@ -70,6 +70,17 @@ function formatCronRunLevelError(error: unknown): string | undefined { return detail ? `cron isolated run failed: ${detail}` : "cron isolated run failed"; } +/** + * Returns the explanation when a reply reports AUTOMATION_FAILED on its exact first line. + * A reply that only quotes the token elsewhere stays ordinary output. + */ +export function readAutomationFailedReport(text: string | undefined): string | undefined { + const [firstLine, ...detail] = (text ?? "").trim().split("\n"); + return firstLine?.trim() === AUTOMATION_FAILED_TOKEN + ? (normalizeOptionalString(detail.join("\n")) ?? "The automation run reported that it failed.") + : undefined; +} + /** Picks a bounded cron run summary from plain text output. */ export function pickSummaryFromOutput(text: string | undefined) { const clean = (text ?? "").trim(); @@ -293,17 +304,9 @@ export function resolveCronPayloadOutcome(params: { ? { ...params.failureSignal, message: failureMessage } : undefined; const runLevelError = formatCronRunLevelError(params.runLevelError); - // Only an exact first line counts; a reply that quotes the token stays ordinary output. - const [reportedFirstLine, ...reportedDetail] = ( - normalizedFinalAssistantVisibleText ?? - fallbackOutputText ?? - "" - ).split("\n"); - const reportedFailure = - reportedFirstLine?.trim() === AUTOMATION_FAILED_TOKEN - ? (normalizeOptionalString(reportedDetail.join("\n")) ?? - "The automation run reported that it failed.") - : undefined; + const reportedFailure = readAutomationFailedReport( + normalizedFinalAssistantVisibleText ?? fallbackOutputText, + ); const hasFatalErrorPayload = hasFatalStructuredErrorPayload || failureSignal !== undefined || diff --git a/src/cron/isolated-agent/run-finalize.ts b/src/cron/isolated-agent/run-finalize.ts index 4f590c828baf..7678cfb517cb 100644 --- a/src/cron/isolated-agent/run-finalize.ts +++ b/src/cron/isolated-agent/run-finalize.ts @@ -227,6 +227,7 @@ export async function finalizeCronRun(params: { outputText, hasFatalErrorPayload, embeddedRunError, + agentReportedFailure, } = cronPayloadOutcome; const terminalToolFailure = finalRunResult.meta?.terminalToolFailure; const hasTerminalToolFailure = isEmbeddedRunTerminalToolFailure(terminalToolFailure); @@ -262,8 +263,8 @@ export async function finalizeCronRun(params: { error: runError, // The agent already judged the task blocked: rerunning it would repeat that turn, // and its prose must not be text-classified into a transient retry reason. - ...(cronPayloadOutcome.agentReportedFailure - ? { errorClassification: { kind: "permanent" as const } } + ...(agentReportedFailure + ? { errorClassification: { kind: "permanent" as const, reportedByAgent: true } } : {}), } : {}), @@ -412,5 +413,10 @@ export async function finalizeCronRun(params: { hasFatalErrorPayload = true; embeddedRunError = pendingPresentationWarningError; } + if (deliveryResult.agentReportedFailure && !hasFatalErrorPayload) { + hasFatalErrorPayload = true; + embeddedRunError = deliveryResult.agentReportedFailure; + agentReportedFailure = true; + } return resolveRunOutcome({ ...deliveryResult, delivery: deliveryTrace }); } diff --git a/src/cron/isolated-agent/run.meta-error-status.test.ts b/src/cron/isolated-agent/run.meta-error-status.test.ts index c726e2291409..52e19c3980f4 100644 --- a/src/cron/isolated-agent/run.meta-error-status.test.ts +++ b/src/cron/isolated-agent/run.meta-error-status.test.ts @@ -168,7 +168,7 @@ describe("runCronIsolatedAgentTurn - meta.error status propagation", () => { expected: { status: "error", error: "Network timeout: no shell tool is available in this run.", - errorClassification: { kind: "permanent" }, + errorClassification: { kind: "permanent", reportedByAgent: true }, }, }, { diff --git a/src/cron/service/timer-outcomes.regression.test.ts b/src/cron/service/timer-outcomes.regression.test.ts index c70a3ebbfd28..b555f56b8faf 100644 --- a/src/cron/service/timer-outcomes.regression.test.ts +++ b/src/cron/service/timer-outcomes.regression.test.ts @@ -174,6 +174,68 @@ describe("cron timer outcome and failure policy regressions", () => { expect(sendCronFailureAlert).not.toHaveBeenCalled(); }); + it.each([ + { name: "silent job", delivery: { mode: "none" as const }, disables: false }, + { + name: "webhook job", + delivery: { mode: "webhook" as const, to: "https://hooks.example.test/cron" }, + disables: true, + }, + { name: "announce job", delivery: undefined, disables: true }, + ])( + "backs off agent-reported failures and auto-disables only with a notification owner: $name", + ({ delivery, disables }) => { + const startedAt = Date.parse("2026-08-01T12:00:00.000Z"); + let now = startedAt; + const deferredNotifications: DeferredCronNotifications = []; + const state = createCronServiceState({ + storePath: "/tmp/cron-reported-failure-threshold.json", + nowMs: () => now, + enqueueSystemEvent: vi.fn(), + runIsolatedAgentJob: createDefaultIsolatedRunner(), + }); + const job = createIsolatedRegressionJob({ + id: "recurring-reported-failure", + name: "recurring reported failure", + scheduledAt: startedAt, + schedule: { kind: "every", everyMs: 60_000, anchorMs: startedAt }, + payload: { kind: "agentTurn", message: "report" }, + state: {}, + }); + job.delivery = delivery; + + const delaysMs: number[] = []; + for (let run = 1; run <= 11 && job.enabled; run += 1) { + applyJobResult( + state, + job, + { + status: "error", + error: "No shell tool is available in this run.", + errorClassification: { kind: "permanent", reportedByAgent: true }, + startedAt: now, + endedAt: now + 10, + }, + { deferredNotifications }, + ); + if (job.state.nextRunAtMs !== undefined) { + delaysMs.push(job.state.nextRunAtMs - (now + 10)); + now = job.state.nextRunAtMs; + } + } + + expect(job.state.lastRunStatus).toBe("error"); + expect(job.state.lastError).toBe("No shell tool is available in this run."); + expect(job.enabled).toBe(!disables); + expect(job.state.consecutiveErrors).toBe(disables ? 10 : 11); + expect(deferredNotifications.filter((n) => n.kind === "auto-disabled")).toHaveLength( + disables ? 1 : 0, + ); + // The streak still drives error backoff past the one-minute cadence: 5 and 15 minutes, then hourly. + expect(delaysMs.slice(2, 6)).toEqual([300_000, 900_000, 3_600_000, 3_600_000]); + }, + ); + it("resets the auto-disable streak after a successful recurring run", () => { const startedAt = Date.parse("2026-08-01T13:00:00.000Z"); const state = createCronServiceState({ diff --git a/src/cron/service/timer-outcomes.ts b/src/cron/service/timer-outcomes.ts index 451c09d5aaa3..34c7dcaed89a 100644 --- a/src/cron/service/timer-outcomes.ts +++ b/src/cron/service/timer-outcomes.ts @@ -1,6 +1,7 @@ import { resolveCronTriggerMinIntervalMs } from "../../config/cron-limits.js"; import type { CronActiveJobMarker } from "../active-jobs.js"; import { resolveAdmittedCronCompletionStatus } from "../completion-status.js"; +import { resolveCronDeliveryPlan } from "../delivery-plan.js"; import { resolvePacedNextRunAtMs } from "../pacing.js"; import { normalizeCronRunDiagnostics, summarizeCronRunDiagnostics } from "../run-diagnostics.js"; import { resolveCronRunErrorReason } from "../run-error-reason.js"; @@ -185,6 +186,14 @@ export function applyJobResult( } }; const alertConfig = resolveFailureAlert(state, job); + // A silent job's agent-reported blocked outcome stays in history, status, and backoff, but no + // notification owner exists for it, so it never auto-disables the job or posts that notice. + const silentReportedFailure = + result.status === "error" && + result.errorClassification?.kind === "permanent" && + result.errorClassification.reportedByAgent === true && + alertConfig === null && + resolveCronDeliveryPlan(job).mode === "none"; if (result.status === "error") { job.state.consecutiveErrors = (job.state.consecutiveErrors ?? 0) + 1; job.state.consecutiveSkipped = 0; @@ -352,6 +361,7 @@ export function applyJobResult( } else if ( result.status === "error" && isJobEnabled(job) && + !silentReportedFailure && maybeAutoDisableCronJobAfterRunFailure({ job, atMs: result.endedAt, diff --git a/src/cron/types.ts b/src/cron/types.ts index 0acde0532b79..fcc897bde904 100644 --- a/src/cron/types.ts +++ b/src/cron/types.ts @@ -151,7 +151,8 @@ export type CronRunDiagnostics = NonNullable /** Explicit execution-error disposition used consistently by retry, history, and alerts. */ export type CronRunErrorClassification = | { kind: "reason"; reason: FailoverReason } - | { kind: "permanent" }; + /** `reportedByAgent`: the run's final answer reported AUTOMATION_FAILED; no runtime fault. */ + | { kind: "permanent"; reportedByAgent?: true }; /** Closed producer-authored facts allowed in operator-facing failure notifications. */ export type CronFailureNotificationDetail =