## Contributor description — original implementation
The contributor’s original description follows. Historical heads, scope and open decisions in this section refer to that implementation; the reviewed current implementation, compatibility decision and proof are recorded in the maintainer update below.
Closes #133245.
Related: #127950. This supersedes the current-main-incompatible queue-owned approach in #125385 while preserving today's canonical timeout retry disposition.
## What Problem This Solves
A durably claimed channel message can wait behind another follow-up turn for longer than the five-minute claim-to-adoption watchdog. Existing liveness renewal covers queue-head active-run admission checks, but not an ingress-backed lifecycle waiting deeper in the follow-up queue. The claim can therefore be retired before the message reaches the model.
## Why This Change Was Made
The follow-up queue now owns periodic deferred heartbeats for the exact ingress lifecycle while it remains in pending items, in-flight delivery, summary sources, or compacted summary elisions.
Ingress supplies a cadence derived from one third of its adoption-stall timeout. Debounce, Plugin SDK fan-in, and channel lifecycle wrappers preserve the shortest applicable cadence. Queue ownership triggers an immediate renewal, then periodic renewal; it stops after successful adoption, completion, ownership loss, or callback failure. A rejected adoption callback keeps renewal alive for the supported retry path.
The existing watchdog and canonical timeout retry policy remain unchanged, so silent handlers and orphaned claims still recover normally.
## User Impact
Messages accepted by durable channel ingress remain adoptable while they wait behind long-running work instead of silently disappearing before model execution. No setting or migration is required.
## Evidence
SDK documentation polish (2026-09-10, head `76097fef9397a04c10637e06dfaf64ebb6ca104a`): the public channel SDK guide now documents forwarding both heartbeat fields, the shortest-positive-finite fan-in cadence, renewal termination, and compatibility when older wrappers omit the optional cadence. The documentation-only addition passed changed-page formatting, MDX sanity, and `git diff --check`; production code and the previously tested runtime head below are unchanged. Current-head hosted CI has completed successfully; the final exact-head ClawSweeper review accepts the implementation and proof, leaving only SDK-owner acceptance of the documented contract.
Refreshed on 2026-09-08 for exact head `57c4faff0ed7bb2d9b4fd308aa46b19ae4095e61`, rebased onto `1ff45cad5a`.
- Preserved upstream's ingress-monitor type extraction and gateway-suspension repair; the optional cadence field follows its new type owner.
- Repaired the retry integration fixture: after real enqueue and initial renewal, only the abandoned message's heartbeat callback fails. The real watchdog then recovers it. Retry delivery, healthy sibling cleanup, and true duplicate assertions remain intact; no production behavior or timeout changed for this repair.
- 102 focused tests pass: dispatch ingress retry, queue in-flight/dedupe, ingress lifecycle/watchdog, Plugin SDK fan-in, Discord queue handling, and Slack handler.
- Changed-file gates pass: core/core-test/extension typechecks, formatting, changed core/extension lint, SDK boundaries/exports, dead-export scans, and repository guards.
- Fresh independent source review found no actionable P0–P2 defects. The standalone autoreview CLI failed at startup without a verdict and is not counted as review coverage.
- Current-head CI run [34520397963](https://github.com/openclaw/openclaw/actions/runs/34520397963) completed successfully. The earlier startup subprocess timeout is historical, not a current failing gate. Exact-head ClawSweeper review (September 10, 20:00 UTC) reports no actionable correctness or proof findings; explicit SDK-owner acceptance remains required before merge.
- Existing regressions cover immediate renewal after late handoff, periodic deeper-queue renewal, stopping after owner loss, and continuing across a rejected adoption callback.
Exact-head durable-ingress boundary proof used isolated SQLite with the production ingress drain/lifecycle binder, follow-up queue, and canonical adoption helper. The model callback was synthetic; this is not live Telegram or Discord transport proof. No production Gateway, channel state, or configuration was touched. Temporary state and queue ownership were cleaned up.
```json
{
"schema": "openclaw.pr133248.durable-ingress-proof.v1",
"exactHead": "57c4faff0ed7bb2d9b4fd308aa46b19ae4095e61",
"setup": {
"durableStore": "isolated OpenClaw SQLite state",
"executor": "production follow-up queue with synthetic model callback",
"transportLifecycle": "production ingress drain lifecycle",
"adoptionStallTimeoutMs": 180
},
"queuedBehindLongTurn": {
"heldMs": 620,
"formerDeadlineCrossed": true,
"heartbeatCount": 11,
"claimStillOwnedAtCheckpoint": [
"queued-event"
],
"retryRowsAtCheckpoint": 0,
"failedRowsAtCheckpoint": 0,
"executionCount": 1,
"duplicateAfterAdoption": "completed"
},
"orphanRecovery": {
"status": "released-for-retry",
"attempts": 1,
"lastErrorContainsHandlerTimeout": true
},
"productionTouched": false
}
```
Owner acceptance:
- Intended behavior: accepted messages remain adoptable while their exact lifecycle is owned by the queue; ownerless claims retain canonical timeout recovery.
- Boundary: one optional cadence field propagated through existing shared lifecycle and channel wrappers; no configuration, schema, or migration change.
- Maintainer decision remains open: accept the additive public Plugin SDK lifecycle field and its queue-owned cadence semantics. Bot review is not that acceptance.
- Rollback: revert the renewal fix and its companion retry-fixture adjustment.
- Scope: 23 files, +273/-2. The width is required lifecycle forwarding; renewal policy stays in the queue and ingress drain.
AI-assisted.
---
## Maintainer update — reviewed head `9325ad501aa4`
## What Problem This Solves
Messages accepted while a long reply is running can expire in the follow-up queue and consume a retry. The reproduction confirmed expiry and retry; all three messages eventually arrived. It did not reproduce the older permanent-loss report in #133245.
## Fix and impact
The queue starts one heartbeat when it accepts a lifecycle and stops it on adoption, completion, cancellation, or callback failure. Heartbeats no longer scan queue collections or copy the in-flight set. Ingress remains responsible for the watchdog and retry settlement, and a heartbeat cannot undo the watchdog pause during adoption finalization.
Ingress derives the cadence from its adoption timeout. The optional lifecycle metadata remains necessary because wrappers rebuild callbacks and combine abort signals, losing the original timeout. No channel setting, schema, migration, dependency, or protocol change is added.
The simplification removes 110 lines net from the initial implementation plus main merge. Total production growth is 40 lines. Existing lifecycle types are reused, and the queue cases share the existing lifecycle test fixtures.
## Evidence
- Real Gateway and Discord: three marked messages on one route, with a queued reply held for 330 seconds. Main expired the third claim after 300.004 seconds. This head retained the same claim at 322.257 seconds with zero attempts; replies arrived once, in order, at 17.804, 349.201, and 350.223 seconds.
- Real SQLite/drain/binder/queue controls: explicit abandonment releases the claim; callback failure stops renewal and lets the watchdog retry without model execution.
- 191 core owner/sibling tests and 190 channel tests passed. All five new regression cases failed on plain main for the intended reasons.
- The local preflight passed core/extension type-aware lint, production and test types, script types, and protocol checks in an isolated checkout of this head.
- External consumers compiled and ran against published 2026.9.4 and this head, including omitted metadata, forwarding, optional return fields, and asynchronous abandonment.
## Consumers
Shared ingress binding, batching, reply dispatch, and the Feishu, Slack, Telegram, and Twitch wrappers preserve the source cadence. The fan-in uses the shortest valid cadence. Legacy wrappers that omit the optional field retain their existing head-only heartbeat behavior. Adoption, queue clearing, overflow, cancellation, summaries, and drain replacement retain or finish the same lifecycle owner.
Closes #133245.
Original implementation and report by @PollyBot13 (#133245). The contributor's commits and authorship are retained.
AI-assisted.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
22 KiB
| summary | title | doc-schema-version | read_when | |||
|---|---|---|---|---|---|---|
| Outbound message lifecycle API for channel plugins: adapters, receipts, durable sends, live preview, and reply pipeline helpers | Channel outbound API | 1 |
|
Channel plugins expose outbound message behavior from
openclaw/plugin-sdk/channel-outbound. Use
openclaw/plugin-sdk/channel-inbound for receive/context/dispatch
orchestration.
Core owns queueing, durability, the durable ingress monitor and drain
(createChannelIngressMonitor, createChannelIngressDrain, and
openChannelIngressDrain), generic retry policy, turn-adoption lifecycle
(turnAdoptionLifecycle / bindIngressLifecycleToReplyOptions), hooks,
receipts, and the shared message tool. The plugin owns native
send/edit/delete calls, target normalization, platform threading, selected
quotes, notification flags, account state, ingress inspection and payload
encoding, lane keys, non-retryable predicates, optional supersede
authorization, and platform-specific side effects.
Durable ingress monitors
Use createChannelIngressMonitor(...) when a channel must persist accepted
transport events before dispatch. It composes a channel ingress queue and drain
with the shared admission, polling, pruning, delivery, and shutdown lifecycle.
Use the lower-level createChannelIngressDrain(...) only when the transport
owns a materially different admission or pump contract.
The required options are:
| Option | Contract |
|---|---|
queue |
A ChannelIngressQueue, or a lazy factory that opens the account-scoped queue. |
inspect(raw, context) |
Returns the stable eventId and serialized laneKey, or null for an ignored event. Claim-time facts must match the persisted id and lane. |
payload |
Supplies the payload version plus body serialization/deserialization. Use storage: "raw-event" for the standard { version, rawEvent } string envelope, or provide custom encode/decode callbacks for an existing channel-specific shape. createClaimError classifies invalid versions or changed identity. |
deliver(raw, lifecycle, claim) |
Dispatches one decoded event and receives the complete adoption lifecycle. It may return completed, deferred, failed-retryable, or nothing. |
pollIntervalMs |
Schedules recovery/drain polls while the monitor is running. |
retention |
Supplies the prune cadence and completed/failed TTL and entry caps. |
The monitor serializes admissions so append backoff cannot invert a lane. The
default bounded append delays are 0, 100, and 300 ms; exhaustion rejects
the transport callback instead of dispatching an event that was not made
durable. At claim time it decodes the versioned payload, re-runs inspect, and
rejects an id or lane mismatch before delivery.
onDurableAdmission(raw, context) runs after every durable enqueue, including
duplicates. context.isNew is true if and only if this admission inserted the
(queue_name, event_id) row. It does not indicate claim ownership or eventual
delivery. If retention previously pruned the row, a later admission may insert
it again and report isNew: true.
deliver receives onAdopted, onDeferred, onAdoptionFinalizing, onFailed,
onCancelled, onAbandoned, and abortSignal. Use onFailed for delivery
errors, onCancelled for explicit pre-adoption cancellation that must preserve
retry accounting, and onAbandoned when a non-adopted turn should consume a
retry attempt. Returning without an explicit handoff marks a terminal
no-dispatch event adopted. admission is always exclusive. A deferred handoff
keeps the claim held, while shutdown or abort leaves unadopted work retryable.
The monitor tracks delivery independently from claim settlement because
adoption can tombstone a row before the channel's delivery promise returns.
One turn, several durable claims
A channel that answers several inbound events as one turn holds one durable
claim per event, and every one of them has to reach a terminal disposition.
fanInChannelIngressLifecycles(lifecycles), from
openclaw/plugin-sdk/channel-ingress-runtime, returns a single
ChannelIngressLifecycle that fans each callback out across all of them, plus
settle, abandon, and cancel for the paths that finish outside a callback.
Pass the events' lifecycles in the order they were claimed; undefined entries
are skipped, and an empty input yields lifecycle: undefined so a caller can
fall back to the single-claim path without branching. Adopting the combined
lifecycle adopts every claim, and abandoning it returns every claim to its retry
budget — no claim is left half-settled. Adoption runs in the order the claims
were passed, and a claim whose adoption is rejected — the drain raises that when
another owner has taken it — fails itself and every claim behind it while the
claims already adopted stay adopted, so a rejection reaches the caller with
nothing left held.
Use it only when one agent turn genuinely consumes several claims. A channel that answers each event on its own keeps passing that event's lifecycle directly.
Start slots and deferral
drain.startLimit bounds how many deliveries the drain starts at once. A
delivery that defers normally keeps its slot, because a deferred delivery on the
default deferredLaneOccupancy: "hold" still serializes its lane. A drain that
declares deferredLaneOccupancy: "release" gives the lane up on deferral, so
its deferred deliveries also give their start slot back, bounded by a budget
equal to startLimit: open delivery callbacks stay within startLimit plus that
budget, and past it a deferral keeps its slot. Without the release, a handful of
deferred deliveries would hold every slot and stall the other lanes. How much
handed-off deferred work may be pending at once stays the drain owner's
semantics, unchanged by this budget.
Optional settings include custom append delays, a drain option block for
advanced drain ordering/concurrency/retry policy, an external abortSignal, a
clock, pump error reporting, a stopped-error factory, and admission policy.
The returned monitor exposes admit, ensureQueueAvailable, start, pause,
stop, waitForIdle, isRunning, and isStopped. Use the idempotent
ensureQueueAvailable() check when plugin-owned migration or preparation must
run after the queue opens but before the drain starts. stop first settles
accepted admissions, then aborts and disposes the drain, waits for the pump and
active deliveries, and disposes again to close the lazy-creation race.
Keep transport-specific redaction, raw-envelope validation, non-retryable
classification, and persisted payload shape in the plugin. Webhook transports
should acknowledge only after admit resolves; non-replay transports should
surface durable append exhaustion rather than silently dispatching.
Deferred claim heartbeats
Forward both onDeferredHeartbeat and deferredHeartbeatIntervalMs when a
plugin wraps the ingress lifecycle or maps it to turnAdoptionLifecycle.
bindIngressLifecycleToReplyOptions(...) forwards both. The drain derives the
optional cadence from its adoption-stall timeout; fan-in uses the shortest
positive, finite source cadence. The queue renews only while it owns the
lifecycle, stopping after adoption, completion, ownership loss, or callback
failure. A heartbeat does not adopt or complete a claim.
Wrappers that omit the cadence remain valid but do not enable periodic renewal; their deferred claims can still reach the adoption watchdog timeout. Plugins must not run independent timers that keep abandoned work alive.
Adapter
Most plugins define one message adapter:
import {
defineChannelMessageAdapter,
createMessageReceiptFromOutboundResults,
} from "openclaw/plugin-sdk/channel-outbound";
export const demoMessageAdapter = defineChannelMessageAdapter({
id: "demo",
durableFinal: {
capabilities: {
text: true,
replyTo: true,
thread: true,
messageSendingHooks: true,
},
},
send: {
text: async ({ cfg, to, text, accountId, replyToId, threadId, signal }) => {
const sent = await sendDemoMessage({
cfg,
to,
text,
accountId: accountId ?? undefined,
replyToId: replyToId ?? undefined,
threadId: threadId == null ? undefined : String(threadId),
signal,
});
return {
receipt: createMessageReceiptFromOutboundResults({
results: [{ channel: "demo", messageId: sent.id, conversationId: to }],
kind: "text",
threadId: threadId == null ? undefined : String(threadId),
replyToId: replyToId ?? undefined,
}),
};
},
},
});
Only declare capabilities the native transport actually preserves. Cover each declared capability with the matching contract helper exported from this subpath:
- send:
verifyChannelMessageAdapterCapabilityProofs(...) - durable final delivery:
verifyDurableFinalCapabilityProofs(...) - live preview:
verifyChannelMessageLiveCapabilityAdapterProofs(...)andverifyChannelMessageLiveFinalizerProofs(...) - receive ack:
verifyChannelMessageReceiveAckPolicyAdapterProofs(...)
Outbound echo suppression
When a platform may redeliver the plugin's own outbound message as inbound, call recordOutboundMessageIdentity(...) with the channel, account, conversation, and a stable platform message or source identity. The shared inbound turn path drops matching identities for a bounded 30-second window before session recording or agent dispatch; a source identity may be reserved before send or refreshed when a channel route is removed to close delivery races. isRecentOutboundMessageIdentity(...) exposes the same query for channel diagnostics and tests. Do not maintain a parallel channel-local TTL cache for the same stable identity.
Plain-text sanitization
Use sanitizeForPlainText(...) when an outbound adapter needs to convert the
supported HTML formatting tags into lightweight text markup. The default keeps
the existing chat-style bold and strikethrough markers. Pass
{ style: "markdown" } only when the channel reparses the result as Markdown:
import { sanitizeForPlainText } from "openclaw/plugin-sdk/channel-outbound";
const chatText = sanitizeForPlainText(text);
const markdownText = sanitizeForPlainText(text, { style: "markdown" });
The Markdown style uses **bold** and ~~strikethrough~~; italic and inline
code keep _italic_ and backtick markers in both styles. Select the style at
the channel boundary instead of rewriting marker text after sanitization.
Delivery Evidence
A MessageReceipt records the result returned by a channel adapter. Concrete
platform message identifiers show that the platform send path accepted the
message; they do not prove that a recipient's device displayed or read it.
Destination and routing identifiers such as chat, channel, room, conversation,
or recipient JID are metadata, never platformMessageIds. Receipts without
platform message identifiers are local receipt metadata only. A
provider-observed receipt thread overrides the requested route thread. If a
batch contains conflicting provider threads, each part retains its thread and
the aggregate receipt omits threadId. Channels with read receipts or
device-delivery state should track those facts through a separate
channel-specific path.
When an adapter intentionally omits a send before dispatch, return
outcome: "not_sent" with an empty receipt and no message ID (legacy outbound
adapters use an empty messageId). Core records adapter_returned_no_send as
an intentional suppression, counts no physical send, and skips send-success
and commit hooks. Do not use this outcome for an acknowledged send without a
platform ID or for an unknown result after dispatch. An empty receipt alone
does not distinguish those states; existing acknowledgement behavior is
unchanged when outcome is omitted.
If a channel adapter can prove that retrying a failure cannot duplicate a
recipient-visible send and no finalization-capable call began, throw
new PlatformMessageNotDispatchedError("...", { cause: error }) from
openclaw/plugin-sdk/error-runtime. Core can then clear stale send-attempt
evidence and safely retry the queued intent. Only the adapter that owns the
final dispatch boundary may make this assertion. Never use the marker after a
finalization/send call begins or returns an ambiguous result; false marking can
duplicate messages.
Existing outbound adapters
If the channel already has a compatible outbound adapter, derive the
message adapter instead of duplicating send code:
import { createChannelMessageAdapterFromOutbound } from "openclaw/plugin-sdk/channel-outbound";
export const messageAdapter = createChannelMessageAdapterFromOutbound({
id: "demo",
outbound,
durableFinal: {
capabilities: {
text: true,
media: true,
},
},
});
Deriving an adapter does not make a channel-owned prepared dispatcher durable.
Route its final sends through the durable helpers while preserving
channel-specific post-send effects and callback-only transport targets.
message.send.lifecycle.afterSendSuccess runs after the native send succeeds;
for queued sends, afterCommit runs after queue acknowledgment. Keep effects
at the boundary they require rather than leaving them only in a legacy dispatcher.
Durable sends
Runtime send helpers also live on channel-outbound:
sendDurableMessageBatch(...)withDurableMessageSendContext(...)deliverInboundReplyWithMessageSendContext(...)- draft streaming/progress helpers such as
resolveChannelDraftStreamingChunking(...)
sendDurableMessageBatch(...) and withDurableMessageSendContext(...) default
to durability: "required": failure to persist the send intent stops delivery
before the platform call. With durability: "best_effort", a queue-write
failure can fall through to a logged, live-only send without crash recovery.
These durable helpers do not accept durability: "disabled".
sendDurableMessageBatch(...) returns one explicit outcome:
| Outcome | Meaning |
|---|---|
sent |
at least one visible platform message was accepted by the platform send path |
suppressed |
no platform message should be treated as missing |
partial_failed |
at least one platform message was accepted before a later payload or side effect failed |
failed |
no platform receipt was produced |
Use payloadOutcomes when a batch mixes sent, suppressed, and failed
payloads. Do not infer hook cancellation from an empty legacy
direct-delivery result.
A suppressed result is ambiguous when its reason is
adapter_returned_no_identity, or any payload outcome records a send, an
identityless send, or a failure with sentBeforeError. Check the whole batch
before treating suppression as intentional non-delivery. For a known omission,
an action can return { status: "suppressed", reason: result.reason } without
a tool error or a fabricated receipt. Keep free-form hook diagnostics private.
Suppression does not recover an earlier failed send.
Failure is not permission to send the same payload through another path.
Once admitted, the queue owns retry or reconciliation until its exact owner
acknowledges or terminally retires the intent. Pending custody is not a
delivery receipt: preserve partial receipts, and do not report an unconfirmed
ask_user prompt as visible. Gateway OUTBOUND_DELIVERY_QUEUED responses mean
delivery is pending and must not be resent; an ambiguous send may need
reconciliation rather than an automatic retry.
When a transport creates a thread during its first successful send, the
outbound adapter may implement adoptTargetFromDelivery(...). Return the
typed thread ID from the platform receipt and core carries it into later
payloads, pins, and post-delivery hooks in that durable batch. Core never
replaces an explicit caller thread, and it does not infer adoption from
receipt.threadId without the adapter opt-in.
Automatic unknown-send reconciliation
Set message.durableFinal.automaticUnknownSendReconciliation only when the
plugin can reconcile an ambiguous provider send from persisted, post-policy
state without rerunning modifying hooks or regenerating provider payloads.
Core considers this opt-in after hooks and cancellation, and only for exactly
one accepted prepared payload. Multi-payload batches do not opt in
automatically.
The adapter must also advertise capabilities.reconcileUnknownSend: true and
provide reconcileUnknownSend(...). Use reconcileUnknownSendKinds to name
the concrete transport branches the plugin can prove, such as text or
media. If the kind map is present, the selected branch must be true.
Omitting the map means the callback claims every selected branch, so prefer an
explicit map for new plugins.
The callback must use provider-owned idempotency or authoritative readback to
return sent with the actual provider receipt, not_sent only when a fresh
send is provably safe, or unresolved when neither outcome can be proven.
When reconciliation is explicitly required, unsupported prepared shapes fail
before provider I/O. During recovery, missing, incomplete, or mismatched
provider proof must fail closed rather than replaying content that could
already be visible.
If reconciliation needs provider-owned persisted evidence, implement
afterUnknownSendTerminal(...). Core calls it after the ambiguous queue row
has authoritatively moved to failed, including retry-budget exhaustion. Use it
to remove provider-owned plans or payloads that are no longer needed. Cleanup
is best effort and must be idempotent; a failure is logged without making the
terminal queue row replayable again.
Deferred delivery admission
Use message.durableFinal.admitDeferredDelivery(...) when a resolved account
cannot safely accept core-managed outbound or deferred delivery. Core calls
this hook synchronously before live outbound work, including paths that skip
queue persistence, and again before replaying a recovered intent. The context
includes cfg, channel, to, accountId, and a phase of live or
recovery.
Return { status: "allowed" } to continue. Return
{ status: "permanent_rejection", reason } when the delivery must not be
persisted, sent directly, or replayed. A live rejection fails before queue
creation, message hooks, or platform work. A recovery rejection marks the
queued record failed and skips reconciliation and replay. Omitting the hook
means allowed.
The hook is a synchronous admission decision, not a send path. Read only
already-loaded config or runtime state; do not perform network, filesystem, or
other asynchronous I/O. Contract tests should exercise both phases and both
result variants through ChannelMessageDurableFinalAdapter from
openclaw/plugin-sdk/channel-outbound.
Compatibility dispatch
Assemble inbound reply dispatch through dispatchChannelInboundReply(...)
from channel-inbound. Keep platform delivery in the delivery adapter; use
channel-outbound for message adapters, durable sends, receipts, live
preview, and reply pipeline options.
Migrating from channel-message
openclaw/plugin-sdk/channel-message is a deprecated compatibility entrypoint.
It still re-exports channel-outbound and preserves three dispatch aliases.
Migrate those aliases to openclaw/plugin-sdk/channel-inbound:
| Deprecated alias | Replacement |
|---|---|
hasFinalChannelTurnDispatch |
hasFinalInboundReplyDispatch |
hasVisibleChannelTurnDispatch |
hasVisibleInboundReplyDispatch |
resolveChannelTurnDispatchCounts |
resolveInboundReplyDispatchCounts |
Follow the dated removal-eligibility window in Migration. This subpath is not tied to the next Plugin SDK major, and eligibility does not itself remove an export. External imports do not emit a runtime warning; update plugin imports rather than waiting for one.
Related
- Channel inbound API — the receive side that records and dispatches before a reply is sent
- Channel ingress API — the resolver that produces the participant identity a send is attributed to
- Building channel plugins — the full channel plugin walkthrough
- Plugin SDK subpaths — which subpath exports each helper
- Plugin SDK migration — removal-eligibility windows for legacy outbound exports