openclaw/docs/plugins/sdk-channel-outbound.md
PollyBot13 dba4da1f8b
fix(channels): keep deeper queued messages alive through long turns (#133248)
## Contributor description — original implementation

The contributor’s original description follows. Historical heads, scope and open decisions in this section refer to that implementation; the reviewed current implementation, compatibility decision and proof are recorded in the maintainer update below.

Closes #133245.

Related: #127950. This supersedes the current-main-incompatible queue-owned approach in #125385 while preserving today's canonical timeout retry disposition.

## What Problem This Solves

A durably claimed channel message can wait behind another follow-up turn for longer than the five-minute claim-to-adoption watchdog. Existing liveness renewal covers queue-head active-run admission checks, but not an ingress-backed lifecycle waiting deeper in the follow-up queue. The claim can therefore be retired before the message reaches the model.

## Why This Change Was Made

The follow-up queue now owns periodic deferred heartbeats for the exact ingress lifecycle while it remains in pending items, in-flight delivery, summary sources, or compacted summary elisions.

Ingress supplies a cadence derived from one third of its adoption-stall timeout. Debounce, Plugin SDK fan-in, and channel lifecycle wrappers preserve the shortest applicable cadence. Queue ownership triggers an immediate renewal, then periodic renewal; it stops after successful adoption, completion, ownership loss, or callback failure. A rejected adoption callback keeps renewal alive for the supported retry path.

The existing watchdog and canonical timeout retry policy remain unchanged, so silent handlers and orphaned claims still recover normally.

## User Impact

Messages accepted by durable channel ingress remain adoptable while they wait behind long-running work instead of silently disappearing before model execution. No setting or migration is required.

## Evidence

SDK documentation polish (2026-09-10, head `76097fef9397a04c10637e06dfaf64ebb6ca104a`): the public channel SDK guide now documents forwarding both heartbeat fields, the shortest-positive-finite fan-in cadence, renewal termination, and compatibility when older wrappers omit the optional cadence. The documentation-only addition passed changed-page formatting, MDX sanity, and `git diff --check`; production code and the previously tested runtime head below are unchanged. Current-head hosted CI has completed successfully; the final exact-head ClawSweeper review accepts the implementation and proof, leaving only SDK-owner acceptance of the documented contract.

Refreshed on 2026-09-08 for exact head `57c4faff0ed7bb2d9b4fd308aa46b19ae4095e61`, rebased onto `1ff45cad5a`.

- Preserved upstream's ingress-monitor type extraction and gateway-suspension repair; the optional cadence field follows its new type owner.
- Repaired the retry integration fixture: after real enqueue and initial renewal, only the abandoned message's heartbeat callback fails. The real watchdog then recovers it. Retry delivery, healthy sibling cleanup, and true duplicate assertions remain intact; no production behavior or timeout changed for this repair.
- 102 focused tests pass: dispatch ingress retry, queue in-flight/dedupe, ingress lifecycle/watchdog, Plugin SDK fan-in, Discord queue handling, and Slack handler.
- Changed-file gates pass: core/core-test/extension typechecks, formatting, changed core/extension lint, SDK boundaries/exports, dead-export scans, and repository guards.
- Fresh independent source review found no actionable P0–P2 defects. The standalone autoreview CLI failed at startup without a verdict and is not counted as review coverage.
- Current-head CI run [34520397963](https://github.com/openclaw/openclaw/actions/runs/34520397963) completed successfully. The earlier startup subprocess timeout is historical, not a current failing gate. Exact-head ClawSweeper review (September 10, 20:00 UTC) reports no actionable correctness or proof findings; explicit SDK-owner acceptance remains required before merge.
- Existing regressions cover immediate renewal after late handoff, periodic deeper-queue renewal, stopping after owner loss, and continuing across a rejected adoption callback.

Exact-head durable-ingress boundary proof used isolated SQLite with the production ingress drain/lifecycle binder, follow-up queue, and canonical adoption helper. The model callback was synthetic; this is not live Telegram or Discord transport proof. No production Gateway, channel state, or configuration was touched. Temporary state and queue ownership were cleaned up.

```json
{
  "schema": "openclaw.pr133248.durable-ingress-proof.v1",
  "exactHead": "57c4faff0ed7bb2d9b4fd308aa46b19ae4095e61",
  "setup": {
    "durableStore": "isolated OpenClaw SQLite state",
    "executor": "production follow-up queue with synthetic model callback",
    "transportLifecycle": "production ingress drain lifecycle",
    "adoptionStallTimeoutMs": 180
  },
  "queuedBehindLongTurn": {
    "heldMs": 620,
    "formerDeadlineCrossed": true,
    "heartbeatCount": 11,
    "claimStillOwnedAtCheckpoint": [
      "queued-event"
    ],
    "retryRowsAtCheckpoint": 0,
    "failedRowsAtCheckpoint": 0,
    "executionCount": 1,
    "duplicateAfterAdoption": "completed"
  },
  "orphanRecovery": {
    "status": "released-for-retry",
    "attempts": 1,
    "lastErrorContainsHandlerTimeout": true
  },
  "productionTouched": false
}
```

Owner acceptance:
- Intended behavior: accepted messages remain adoptable while their exact lifecycle is owned by the queue; ownerless claims retain canonical timeout recovery.
- Boundary: one optional cadence field propagated through existing shared lifecycle and channel wrappers; no configuration, schema, or migration change.
- Maintainer decision remains open: accept the additive public Plugin SDK lifecycle field and its queue-owned cadence semantics. Bot review is not that acceptance.
- Rollback: revert the renewal fix and its companion retry-fixture adjustment.
- Scope: 23 files, +273/-2. The width is required lifecycle forwarding; renewal policy stays in the queue and ingress drain.

AI-assisted.

---

## Maintainer update — reviewed head `9325ad501aa4`

## What Problem This Solves

Messages accepted while a long reply is running can expire in the follow-up queue and consume a retry. The reproduction confirmed expiry and retry; all three messages eventually arrived. It did not reproduce the older permanent-loss report in #133245.

## Fix and impact

The queue starts one heartbeat when it accepts a lifecycle and stops it on adoption, completion, cancellation, or callback failure. Heartbeats no longer scan queue collections or copy the in-flight set. Ingress remains responsible for the watchdog and retry settlement, and a heartbeat cannot undo the watchdog pause during adoption finalization.

Ingress derives the cadence from its adoption timeout. The optional lifecycle metadata remains necessary because wrappers rebuild callbacks and combine abort signals, losing the original timeout. No channel setting, schema, migration, dependency, or protocol change is added.

The simplification removes 110 lines net from the initial implementation plus main merge. Total production growth is 40 lines. Existing lifecycle types are reused, and the queue cases share the existing lifecycle test fixtures.

## Evidence

- Real Gateway and Discord: three marked messages on one route, with a queued reply held for 330 seconds. Main expired the third claim after 300.004 seconds. This head retained the same claim at 322.257 seconds with zero attempts; replies arrived once, in order, at 17.804, 349.201, and 350.223 seconds.
- Real SQLite/drain/binder/queue controls: explicit abandonment releases the claim; callback failure stops renewal and lets the watchdog retry without model execution.
- 191 core owner/sibling tests and 190 channel tests passed. All five new regression cases failed on plain main for the intended reasons.
- The local preflight passed core/extension type-aware lint, production and test types, script types, and protocol checks in an isolated checkout of this head.
- External consumers compiled and ran against published 2026.9.4 and this head, including omitted metadata, forwarding, optional return fields, and asynchronous abandonment.

## Consumers

Shared ingress binding, batching, reply dispatch, and the Feishu, Slack, Telegram, and Twitch wrappers preserve the source cadence. The fan-in uses the shortest valid cadence. Legacy wrappers that omit the optional field retain their existing head-only heartbeat behavior. Adoption, queue clearing, overflow, cancellation, summaries, and drain replacement retain or finish the same lifecycle owner.

Closes #133245.

Original implementation and report by @PollyBot13 (#133245). The contributor's commits and authorship are retained.

AI-assisted.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-14 10:52:52 +05:30

22 KiB

summary title doc-schema-version read_when
Outbound message lifecycle API for channel plugins: adapters, receipts, durable sends, live preview, and reply pipeline helpers Channel outbound API 1
You are building or refactoring a messaging channel plugin send path
You need durable final reply delivery, receipts, live preview finalization, or receive acknowledgement policy
You are migrating from channel-message or legacy reply dispatch helpers

Channel plugins expose outbound message behavior from openclaw/plugin-sdk/channel-outbound. Use openclaw/plugin-sdk/channel-inbound for receive/context/dispatch orchestration.

Core owns queueing, durability, the durable ingress monitor and drain (createChannelIngressMonitor, createChannelIngressDrain, and openChannelIngressDrain), generic retry policy, turn-adoption lifecycle (turnAdoptionLifecycle / bindIngressLifecycleToReplyOptions), hooks, receipts, and the shared message tool. The plugin owns native send/edit/delete calls, target normalization, platform threading, selected quotes, notification flags, account state, ingress inspection and payload encoding, lane keys, non-retryable predicates, optional supersede authorization, and platform-specific side effects.

Durable ingress monitors

Use createChannelIngressMonitor(...) when a channel must persist accepted transport events before dispatch. It composes a channel ingress queue and drain with the shared admission, polling, pruning, delivery, and shutdown lifecycle. Use the lower-level createChannelIngressDrain(...) only when the transport owns a materially different admission or pump contract.

The required options are:

Option Contract
queue A ChannelIngressQueue, or a lazy factory that opens the account-scoped queue.
inspect(raw, context) Returns the stable eventId and serialized laneKey, or null for an ignored event. Claim-time facts must match the persisted id and lane.
payload Supplies the payload version plus body serialization/deserialization. Use storage: "raw-event" for the standard { version, rawEvent } string envelope, or provide custom encode/decode callbacks for an existing channel-specific shape. createClaimError classifies invalid versions or changed identity.
deliver(raw, lifecycle, claim) Dispatches one decoded event and receives the complete adoption lifecycle. It may return completed, deferred, failed-retryable, or nothing.
pollIntervalMs Schedules recovery/drain polls while the monitor is running.
retention Supplies the prune cadence and completed/failed TTL and entry caps.

The monitor serializes admissions so append backoff cannot invert a lane. The default bounded append delays are 0, 100, and 300 ms; exhaustion rejects the transport callback instead of dispatching an event that was not made durable. At claim time it decodes the versioned payload, re-runs inspect, and rejects an id or lane mismatch before delivery.

onDurableAdmission(raw, context) runs after every durable enqueue, including duplicates. context.isNew is true if and only if this admission inserted the (queue_name, event_id) row. It does not indicate claim ownership or eventual delivery. If retention previously pruned the row, a later admission may insert it again and report isNew: true.

deliver receives onAdopted, onDeferred, onAdoptionFinalizing, onFailed, onCancelled, onAbandoned, and abortSignal. Use onFailed for delivery errors, onCancelled for explicit pre-adoption cancellation that must preserve retry accounting, and onAbandoned when a non-adopted turn should consume a retry attempt. Returning without an explicit handoff marks a terminal no-dispatch event adopted. admission is always exclusive. A deferred handoff keeps the claim held, while shutdown or abort leaves unadopted work retryable. The monitor tracks delivery independently from claim settlement because adoption can tombstone a row before the channel's delivery promise returns.

One turn, several durable claims

A channel that answers several inbound events as one turn holds one durable claim per event, and every one of them has to reach a terminal disposition. fanInChannelIngressLifecycles(lifecycles), from openclaw/plugin-sdk/channel-ingress-runtime, returns a single ChannelIngressLifecycle that fans each callback out across all of them, plus settle, abandon, and cancel for the paths that finish outside a callback. Pass the events' lifecycles in the order they were claimed; undefined entries are skipped, and an empty input yields lifecycle: undefined so a caller can fall back to the single-claim path without branching. Adopting the combined lifecycle adopts every claim, and abandoning it returns every claim to its retry budget — no claim is left half-settled. Adoption runs in the order the claims were passed, and a claim whose adoption is rejected — the drain raises that when another owner has taken it — fails itself and every claim behind it while the claims already adopted stay adopted, so a rejection reaches the caller with nothing left held.

Use it only when one agent turn genuinely consumes several claims. A channel that answers each event on its own keeps passing that event's lifecycle directly.

Start slots and deferral

drain.startLimit bounds how many deliveries the drain starts at once. A delivery that defers normally keeps its slot, because a deferred delivery on the default deferredLaneOccupancy: "hold" still serializes its lane. A drain that declares deferredLaneOccupancy: "release" gives the lane up on deferral, so its deferred deliveries also give their start slot back, bounded by a budget equal to startLimit: open delivery callbacks stay within startLimit plus that budget, and past it a deferral keeps its slot. Without the release, a handful of deferred deliveries would hold every slot and stall the other lanes. How much handed-off deferred work may be pending at once stays the drain owner's semantics, unchanged by this budget.

Optional settings include custom append delays, a drain option block for advanced drain ordering/concurrency/retry policy, an external abortSignal, a clock, pump error reporting, a stopped-error factory, and admission policy. The returned monitor exposes admit, ensureQueueAvailable, start, pause, stop, waitForIdle, isRunning, and isStopped. Use the idempotent ensureQueueAvailable() check when plugin-owned migration or preparation must run after the queue opens but before the drain starts. stop first settles accepted admissions, then aborts and disposes the drain, waits for the pump and active deliveries, and disposes again to close the lazy-creation race.

Keep transport-specific redaction, raw-envelope validation, non-retryable classification, and persisted payload shape in the plugin. Webhook transports should acknowledge only after admit resolves; non-replay transports should surface durable append exhaustion rather than silently dispatching.

Deferred claim heartbeats

Forward both onDeferredHeartbeat and deferredHeartbeatIntervalMs when a plugin wraps the ingress lifecycle or maps it to turnAdoptionLifecycle. bindIngressLifecycleToReplyOptions(...) forwards both. The drain derives the optional cadence from its adoption-stall timeout; fan-in uses the shortest positive, finite source cadence. The queue renews only while it owns the lifecycle, stopping after adoption, completion, ownership loss, or callback failure. A heartbeat does not adopt or complete a claim.

Wrappers that omit the cadence remain valid but do not enable periodic renewal; their deferred claims can still reach the adoption watchdog timeout. Plugins must not run independent timers that keep abandoned work alive.

Adapter

Most plugins define one message adapter:

import {
  defineChannelMessageAdapter,
  createMessageReceiptFromOutboundResults,
} from "openclaw/plugin-sdk/channel-outbound";

export const demoMessageAdapter = defineChannelMessageAdapter({
  id: "demo",
  durableFinal: {
    capabilities: {
      text: true,
      replyTo: true,
      thread: true,
      messageSendingHooks: true,
    },
  },
  send: {
    text: async ({ cfg, to, text, accountId, replyToId, threadId, signal }) => {
      const sent = await sendDemoMessage({
        cfg,
        to,
        text,
        accountId: accountId ?? undefined,
        replyToId: replyToId ?? undefined,
        threadId: threadId == null ? undefined : String(threadId),
        signal,
      });

      return {
        receipt: createMessageReceiptFromOutboundResults({
          results: [{ channel: "demo", messageId: sent.id, conversationId: to }],
          kind: "text",
          threadId: threadId == null ? undefined : String(threadId),
          replyToId: replyToId ?? undefined,
        }),
      };
    },
  },
});

Only declare capabilities the native transport actually preserves. Cover each declared capability with the matching contract helper exported from this subpath:

  • send: verifyChannelMessageAdapterCapabilityProofs(...)
  • durable final delivery: verifyDurableFinalCapabilityProofs(...)
  • live preview: verifyChannelMessageLiveCapabilityAdapterProofs(...) and verifyChannelMessageLiveFinalizerProofs(...)
  • receive ack: verifyChannelMessageReceiveAckPolicyAdapterProofs(...)

Outbound echo suppression

When a platform may redeliver the plugin's own outbound message as inbound, call recordOutboundMessageIdentity(...) with the channel, account, conversation, and a stable platform message or source identity. The shared inbound turn path drops matching identities for a bounded 30-second window before session recording or agent dispatch; a source identity may be reserved before send or refreshed when a channel route is removed to close delivery races. isRecentOutboundMessageIdentity(...) exposes the same query for channel diagnostics and tests. Do not maintain a parallel channel-local TTL cache for the same stable identity.

Plain-text sanitization

Use sanitizeForPlainText(...) when an outbound adapter needs to convert the supported HTML formatting tags into lightweight text markup. The default keeps the existing chat-style bold and strikethrough markers. Pass { style: "markdown" } only when the channel reparses the result as Markdown:

import { sanitizeForPlainText } from "openclaw/plugin-sdk/channel-outbound";

const chatText = sanitizeForPlainText(text);
const markdownText = sanitizeForPlainText(text, { style: "markdown" });

The Markdown style uses **bold** and ~~strikethrough~~; italic and inline code keep _italic_ and backtick markers in both styles. Select the style at the channel boundary instead of rewriting marker text after sanitization.

Delivery Evidence

A MessageReceipt records the result returned by a channel adapter. Concrete platform message identifiers show that the platform send path accepted the message; they do not prove that a recipient's device displayed or read it. Destination and routing identifiers such as chat, channel, room, conversation, or recipient JID are metadata, never platformMessageIds. Receipts without platform message identifiers are local receipt metadata only. A provider-observed receipt thread overrides the requested route thread. If a batch contains conflicting provider threads, each part retains its thread and the aggregate receipt omits threadId. Channels with read receipts or device-delivery state should track those facts through a separate channel-specific path.

When an adapter intentionally omits a send before dispatch, return outcome: "not_sent" with an empty receipt and no message ID (legacy outbound adapters use an empty messageId). Core records adapter_returned_no_send as an intentional suppression, counts no physical send, and skips send-success and commit hooks. Do not use this outcome for an acknowledged send without a platform ID or for an unknown result after dispatch. An empty receipt alone does not distinguish those states; existing acknowledgement behavior is unchanged when outcome is omitted.

If a channel adapter can prove that retrying a failure cannot duplicate a recipient-visible send and no finalization-capable call began, throw new PlatformMessageNotDispatchedError("...", { cause: error }) from openclaw/plugin-sdk/error-runtime. Core can then clear stale send-attempt evidence and safely retry the queued intent. Only the adapter that owns the final dispatch boundary may make this assertion. Never use the marker after a finalization/send call begins or returns an ambiguous result; false marking can duplicate messages.

Existing outbound adapters

If the channel already has a compatible outbound adapter, derive the message adapter instead of duplicating send code:

import { createChannelMessageAdapterFromOutbound } from "openclaw/plugin-sdk/channel-outbound";

export const messageAdapter = createChannelMessageAdapterFromOutbound({
  id: "demo",
  outbound,
  durableFinal: {
    capabilities: {
      text: true,
      media: true,
    },
  },
});

Deriving an adapter does not make a channel-owned prepared dispatcher durable. Route its final sends through the durable helpers while preserving channel-specific post-send effects and callback-only transport targets. message.send.lifecycle.afterSendSuccess runs after the native send succeeds; for queued sends, afterCommit runs after queue acknowledgment. Keep effects at the boundary they require rather than leaving them only in a legacy dispatcher.

Durable sends

Runtime send helpers also live on channel-outbound:

  • sendDurableMessageBatch(...)
  • withDurableMessageSendContext(...)
  • deliverInboundReplyWithMessageSendContext(...)
  • draft streaming/progress helpers such as resolveChannelDraftStreamingChunking(...)

sendDurableMessageBatch(...) and withDurableMessageSendContext(...) default to durability: "required": failure to persist the send intent stops delivery before the platform call. With durability: "best_effort", a queue-write failure can fall through to a logged, live-only send without crash recovery. These durable helpers do not accept durability: "disabled".

sendDurableMessageBatch(...) returns one explicit outcome:

Outcome Meaning
sent at least one visible platform message was accepted by the platform send path
suppressed no platform message should be treated as missing
partial_failed at least one platform message was accepted before a later payload or side effect failed
failed no platform receipt was produced

Use payloadOutcomes when a batch mixes sent, suppressed, and failed payloads. Do not infer hook cancellation from an empty legacy direct-delivery result.

A suppressed result is ambiguous when its reason is adapter_returned_no_identity, or any payload outcome records a send, an identityless send, or a failure with sentBeforeError. Check the whole batch before treating suppression as intentional non-delivery. For a known omission, an action can return { status: "suppressed", reason: result.reason } without a tool error or a fabricated receipt. Keep free-form hook diagnostics private. Suppression does not recover an earlier failed send.

Failure is not permission to send the same payload through another path. Once admitted, the queue owns retry or reconciliation until its exact owner acknowledges or terminally retires the intent. Pending custody is not a delivery receipt: preserve partial receipts, and do not report an unconfirmed ask_user prompt as visible. Gateway OUTBOUND_DELIVERY_QUEUED responses mean delivery is pending and must not be resent; an ambiguous send may need reconciliation rather than an automatic retry.

When a transport creates a thread during its first successful send, the outbound adapter may implement adoptTargetFromDelivery(...). Return the typed thread ID from the platform receipt and core carries it into later payloads, pins, and post-delivery hooks in that durable batch. Core never replaces an explicit caller thread, and it does not infer adoption from receipt.threadId without the adapter opt-in.

Automatic unknown-send reconciliation

Set message.durableFinal.automaticUnknownSendReconciliation only when the plugin can reconcile an ambiguous provider send from persisted, post-policy state without rerunning modifying hooks or regenerating provider payloads. Core considers this opt-in after hooks and cancellation, and only for exactly one accepted prepared payload. Multi-payload batches do not opt in automatically.

The adapter must also advertise capabilities.reconcileUnknownSend: true and provide reconcileUnknownSend(...). Use reconcileUnknownSendKinds to name the concrete transport branches the plugin can prove, such as text or media. If the kind map is present, the selected branch must be true. Omitting the map means the callback claims every selected branch, so prefer an explicit map for new plugins.

The callback must use provider-owned idempotency or authoritative readback to return sent with the actual provider receipt, not_sent only when a fresh send is provably safe, or unresolved when neither outcome can be proven. When reconciliation is explicitly required, unsupported prepared shapes fail before provider I/O. During recovery, missing, incomplete, or mismatched provider proof must fail closed rather than replaying content that could already be visible.

If reconciliation needs provider-owned persisted evidence, implement afterUnknownSendTerminal(...). Core calls it after the ambiguous queue row has authoritatively moved to failed, including retry-budget exhaustion. Use it to remove provider-owned plans or payloads that are no longer needed. Cleanup is best effort and must be idempotent; a failure is logged without making the terminal queue row replayable again.

Deferred delivery admission

Use message.durableFinal.admitDeferredDelivery(...) when a resolved account cannot safely accept core-managed outbound or deferred delivery. Core calls this hook synchronously before live outbound work, including paths that skip queue persistence, and again before replaying a recovered intent. The context includes cfg, channel, to, accountId, and a phase of live or recovery.

Return { status: "allowed" } to continue. Return { status: "permanent_rejection", reason } when the delivery must not be persisted, sent directly, or replayed. A live rejection fails before queue creation, message hooks, or platform work. A recovery rejection marks the queued record failed and skips reconciliation and replay. Omitting the hook means allowed.

The hook is a synchronous admission decision, not a send path. Read only already-loaded config or runtime state; do not perform network, filesystem, or other asynchronous I/O. Contract tests should exercise both phases and both result variants through ChannelMessageDurableFinalAdapter from openclaw/plugin-sdk/channel-outbound.

Compatibility dispatch

Assemble inbound reply dispatch through dispatchChannelInboundReply(...) from channel-inbound. Keep platform delivery in the delivery adapter; use channel-outbound for message adapters, durable sends, receipts, live preview, and reply pipeline options.

Migrating from channel-message

openclaw/plugin-sdk/channel-message is a deprecated compatibility entrypoint. It still re-exports channel-outbound and preserves three dispatch aliases. Migrate those aliases to openclaw/plugin-sdk/channel-inbound:

Deprecated alias Replacement
hasFinalChannelTurnDispatch hasFinalInboundReplyDispatch
hasVisibleChannelTurnDispatch hasVisibleInboundReplyDispatch
resolveChannelTurnDispatchCounts resolveInboundReplyDispatchCounts

Follow the dated removal-eligibility window in Migration. This subpath is not tied to the next Plugin SDK major, and eligibility does not itself remove an export. External imports do not emit a runtime warning; update plugin imports rather than waiting for one.