* fix(agents): scope subagent concurrency to spawning sessions Give each immediate spawning session its own configured execution budget, including nested orchestrators, instead of making unrelated sessions share one Gateway-wide queue. Preserve queue ordering, cancellation, hot limit updates, and idle cleanup; retain the same parent identity during compaction. Show aggregate activity with an explicit per-session limit in diagnostics. Keep the existing child admission cap and Codex-native scheduling separate. Verify with focused queue/runtime/UI tests and an isolated real-provider Gateway case that holds one session's slot while another session and its nested child run, then awaits all completion acknowledgments. Closes #156370 * test(telegram): match sticker fixture to ingress dispatch Register the native sticker pipeline's provider runtime prerequisite for local and CI execution. Preserve the same Request JSON decoder while exercising the registered bot handler, as durable webhook ingress does after acknowledgment, and join the held turn before fixture cleanup. The original CI file order reproduced the grammY timeout. Runtime preparation alone was insufficient; the test's generic webhook adapter imposed a deadline that is absent from the production dispatch boundary. Keep the cache, media, and model-admission assertions unchanged, with the real webhook acknowledgment contract covered by its existing test. * test: repair packaging and runtime-owner CI fixtures Copy the newly imported check-limits helper into the trusted packaging harness so its dependency-free startup assertion reaches the CLI. Compare prepared Telegram runtime files against the same config that supplies the expected files. Keep worker-envelope coverage, packing, concurrency and execution-budget assertions unchanged. Both regressions reproduced before their fixes. The packaging owner passed 64 tests and ten targeted packing/prerequisite checks passed afterward. * test(matrix): await follow-up adoption while active turn is held Replace the 500 ms post-routing poll with resolver-owned adoption completion. Preserve foreground FIFO settlement by releasing the active turn before joining handlers, and settle fixture gates on timeout. Validation: all four owner cases pass (37.31 s wrapper, 3.32 s bodies); selected changed checks and independent P0-P2 review pass. * test(ci): recognize stopped Linux fixture process groups Reuse the existing all-thread process-group assertion after detached fixture leaders are joined. Kernel kill probes also include exited zombies awaiting reaping; retain rejection of live descendants and uncertain observations. Validation: 64 Docker scheduler cases and 18 retained-zombie/census cases passed on Linux Testbox; selected changed checks and independent P0-P2 review pass. The original CI process state did not reproduce in the original-order diagnostic. * test(ui): count roster refreshes independently of child lookups Install the browser clock before navigation and match roster sessions.list requests by includeGlobal. Preserve the 4999/+2 ms event window and avatar checks, and require exactly one roster refresh after the boundary. Validation: reproduced the failing count and traced its extra request to a spawnedBy child lookup while the roster timer remained pending. All 14 browser cases pass (38.15 s; changed case 1.063 s). Selected changed gates and independent P0-P2 review pass.
15 KiB
| summary | read_when | title | |||
|---|---|---|---|---|---|
| Auto-reply queue modes, shared background capacity, and per-session overrides |
|
Command queue |
OpenClaw serializes inbound auto-reply runs (all channels) through a tiny in-process queue to prevent multiple agent runs from colliding, while still allowing safe parallelism across sessions.
Why
- Auto-reply runs can be expensive (LLM calls) and can collide when multiple inbound messages arrive close together.
- Serializing avoids competing for shared resources (session state, logs, CLI stdin) and reduces the chance of upstream rate limits.
How it works
- A lane-aware FIFO queue drains each lane with a configurable concurrency cap (default 1 for unconfigured lanes;
mainusesmax(8, available CPU parallelism * 4), and sub-agent queues default to 8 per spawning session). - CLI, embedded, and Codex runs share the same session-key lane (
session:<key>). Each turn waits there before acquiring the session's execution claim, so changing runtimes cannot start a competing turn. - Inbound session runs then enter the global
mainlane, whose parallelism is capped byagents.defaults.maxConcurrent. Sub-agent runs instead use their immediate spawning/controller session's budget, set byagents.defaults.subagents.maxConcurrent. - Embedded attempt preparation starts one stage per event-loop turn so concurrent starts leave room for Gateway requests. Asynchronous stage work can still overlap; this does not lower the run concurrency limit or change session serialization.
- When verbose logging is enabled, queued runs emit a short notice if they waited more than ~2s before starting.
- Typing indicators still fire immediately on enqueue (when supported by the channel) so user experience is unchanged while the run waits its turn.
Defaults
When unset, all inbound channel surfaces use:
mode: "steer"- a built-in 500ms debounce for steer, followup, and collect batching
cap: 20drop: "summarize"
Same-turn steering is the default. A prompt that arrives mid-run is injected into the active runtime when the run can accept steering, so no second session run is started. If the active run cannot accept steering, OpenClaw waits for the active run to finish before starting the prompt.
Queue modes
/queue controls what normal inbound messages do while a session already has an active run:
steer: inject messages into the active runtime. OpenClaw lets an already-running tool finish, skips sequential calls that have not started, and makes the steer visible before the next tool launch or model decision. Parallel calls continue once their batch has crossed its launch checkpoint. Codex app-server receives one batchedturn/steerand applies it at the next model boundary. If the run is not actively streaming or steering is unavailable, OpenClaw waits until the active run ends before starting the prompt.followup: do not steer. Enqueue each message for a later agent turn after the current run ends.collect: do not steer. Coalesce queued messages into a single followup turn after the quiet window. If messages target different channels/threads, they drain individually to preserve routing.interrupt: abort the active run for that session, then run the newest message.
For runtime-specific timing and dependency behavior, see Steering queue. For the explicit /steer <message> command, see Steer.
Gateway input retains its authenticated operator and original scope ceiling while queued or delegated to a child. Collecting messages or steering an active run requires compatible operator sources and tool permissions; other input waits in FIFO order instead of borrowing the active or newest sender's permissions.
An accepted turn can continue after its request returns or its client disconnects. That does not extend revoked device authority or permissions removed by a current operator role. Subsequent actions still check the original source, including work held by an accepted child.
Configure globally or per channel via messages.queue:
{
messages: {
queue: {
mode: "steer",
cap: 20,
drop: "summarize",
byChannel: { discord: "collect" },
debounceMsByChannel: { discord: 1000 },
},
},
}
Queue options
Per-session /queue options apply to queued delivery. The debounce option also sets the Codex steering quiet window in steer mode:
debounce: quiet window before draining queued followups or collect batches; in Codexsteermode, quiet window before sending batchedturn/steer. Bare numbers are milliseconds; unitsms,s,m,h, anddare accepted.cap: max queued messages per session. Values below1are ignored.drop: "summarize"(default): drop the oldest queued entries as needed, keep compact summaries, and inject them as a synthetic followup prompt.drop: "old": drop the oldest queued entries as needed, without preserving summaries.drop: "new": reject the newest message when the queue is already full.
The queue uses a built-in 500ms debounce. cap defaults to 20, and drop defaults to summarize.
Steer and streaming
When channel streaming is partial or block, steering can look like several short visible replies while the active run reaches runtime boundaries:
partial: the preview may finalize early, then a new preview starts after steering is accepted.block: draft-sized blocks can create the same sequential appearance.- Without streaming, steering falls back to a followup after the active run when the runtime cannot accept same-turn steering.
steer does not abort in-flight tools. Skipped OpenClaw tool calls receive synthetic paired error results so the transcript remains valid. Use /queue interrupt when the newest message should abort the current run.
Answering a pending question
A plain-text answer to a pending agent question goes to that question before ordinary queue handling, including when a native CLI cannot accept steering. OpenClaw checks the answer against the question creator's permissions and active run, not the model selected for your next turn. Changed permissions or a closed creator produce an explicit refusal rather than starting another turn.
If the answer may have committed but confirmation is lost, OpenClaw reports that uncertainty and does not resend it as steering or a followup. Check the conversation before retrying. A later delivery or source-cleanup failure does not make the answer replayable, and uncertainty alone does not cancel the original agent run.
Precedence
For mode selection, OpenClaw resolves:
- Inline or stored per-session
/queueoverride. messages.queue.byChannel.<channel>.messages.queue.mode.- Default
steer.
For options, inline or stored /queue options win over config. Then channel-specific debounce (messages.queue.debounceMsByChannel), plugin debounce defaults, and built-in defaults are applied, in that order. cap and drop are global/session options, not per-channel config keys.
Per-session overrides
- Send
/queue <steer|followup|collect|interrupt>as a standalone command to store the queue mode for the current session. - Options can be combined:
/queue collect debounce:0.5s cap:25 drop:summarize /queue defaultor/queue resetclears the session override.
Queued-turn cancellation
While a prompt sits in the followup/collect queue (for example a TUI or
webchat chat.send arriving while another turn is active), Gateway keeps a
Gateway-owned cancel identity for that client runId until the queued
content runs or is dropped. The identity follows content folded into an
overflow summary.
chat.abortwith a specificrunIdcancels that turn while it is still queued, if the requester is authorized (same ownership rules as active runs).chat.abortfor a session withoutrunIdcancels authorized queued turns first, then aborts authorized active runs. That order prevents queue drain from promoting work into a half-stopped session.- Clearing the entire session queue without per-requester checks is not the stop path for multi-owner sessions.
- Queued waits are not projected as active agent runs for
sessions.listand do not own active-run timeout semantics; only the active phase does.
Gateway-backed clients (including openclaw tui) forward mid-run prompts and
let the Gateway apply the queue mode. Esc//stop uses a session-scoped abort
so lost local handles cannot leave a still-queued prompt running.
openclaw chat and openclaw tui --local apply the same four modes in the
embedded runtime. Local steer injects into an active embedded run when that
runtime accepts steering and otherwise becomes a followup; followup and
collect remain local pending work; interrupt aborts the active local run
before starting the newest message. The explicit /steer <message> command is
not a local-mode command.
Input durability
Ordinary user input sent through chat.send to an existing session is stored in the per-agent database
before the Gateway acknowledges it. This includes the Control UI, TUI, CLI, native apps, and RPC clients.
Other connected clients can display the accepted input while it waits, without waiting for a new agent turn.
In collect mode, appending the combined
turn and marking its source inputs consumed happen in one transaction. A browser
reconnect can reconcile those source inputs even if it missed their final events.
The chat displays recorded non-Web client sources separately from the sender, for example Alice · via CLI. Web sources are omitted from these labels, including in collected messages that also contain input from another client.
Reported app names describe the submitting client; they do not establish a human identity or grant permissions.
Collected messages retain their contributing client sources, and older messages without recorded sources keep their existing attribution.
This preserves input, not execution permissions. If the Gateway stops before a queued input reaches the transcript, it appears as interrupted input after restart and requires an explicit resend. The in-memory queue is not replayed. Host sleep that preserves the process can continue the existing queue normally.
Channel messages retained by durable ingress remain retryable when a queued attempt is abandoned before agent-turn adoption. Abandonment releases that attempt's inbound and queue dedupe entries before ingress retries it. Messages already adopted or consumed keep duplicate suppression, so transport redelivery does not repeat their effects.
Lanes and scope
- Applies to auto-reply agent runs across all inbound channels that use the gateway reply pipeline (WhatsApp web, Telegram, Slack, Discord, Signal, iMessage, webchat, etc.).
- Default lane (
main) is process-wide for inbound turns; setagents.defaults.maxConcurrentto allow multiple sessions in parallel. - Heartbeat embedded runs use the bounded
cron-nestedlane for global admission so slow background work does not block inbound replies, while their configured heartbeat session lane still serializes work for that session. - Additional lanes may exist (e.g.
cron,cron-nested,nested) so background jobs can run in parallel without blocking inbound replies. Isolated cron agent turns hold acronslot while their inner agent execution usescron-nested. Shared non-cronnestedflows keep their own lane behavior. These detached runs are tracked as background tasks. - Sub-agent execution uses a separate queue per immediate spawning/controller session.
agents.defaults.subagents.maxConcurrentdefaults to8for each session; independent sessions and nested orchestrators do not share those slots. The separatemaxChildrenPerAgentadmission limit still applies. Codex-native subagents use Codex's own scheduler. - Per-session lanes guarantee that only one agent run touches a given session at a time.
- No external dependencies or background worker threads; pure TypeScript + promises.
Background work
Skill Workshop reviews and plugin background completions, including dreaming, share a separate budget of three concurrent runs. Workshop reviews use at most one slot; each plugin can use up to three available slots. This keeps maintenance work out of foreground reply capacity while bounding its total concurrency. These limits are built in and need no configuration.
Schedulers that await background work do not occupy this budget themselves. Only the dispatched work holds a slot, through completion or cancellation cleanup, so a scheduler cannot block the child it is waiting for. Cancelled queued work is removed before it starts; Gateway restart or runtime retirement prevents stale completions from starting or returning results.
The Control UI System busyness overlay and diagnostics.lanes report this work in one background row. Its active and queued counts include every owner; owner lanes are not counted again in the dynamic session-lane totals.
Troubleshooting
- If commands seem stuck, enable verbose logs and look for "queued for ...ms" lines to confirm the queue is draining.
- Codex app-server runs that accept a turn and then stop emitting progress are interrupted by the Codex adapter so the active session lane can release instead of waiting for the outer run timeout.
- When diagnostics are enabled, sessions that remain in
processingpast the built-in warning threshold with no observed reply, tool, status, block, or ACP progress are classified by current activity:- Active work with recent progress logs as
session.long_running. Owned silent model calls also staysession.long_runninguntil the built-in abort threshold so slow or non-streaming providers are not reported as stalled too early. - Active work with no recent progress logs as
session.stalled; owned model calls, blocked tool calls, and stalled embedded runs switch tosession.stalledat or after the abort threshold. Ownerless stale model/tool activity is not hidden as long-running. session.stuckis reserved for recoverable stale session bookkeeping, including idle queued sessions with stale ownerless model/tool activity.session.stuckalways triggers recovery that can release the affected session lane. Asession.stalledclassification past the abort threshold (blocked tool call, stalled model call, or stalled embedded run) can also trigger active-abort recovery, so both classifications can unstick a queue, not onlysession.stuck.- Repeated
session.stuckandsession.long_runningwarning log lines back off exponentially while the session remains unchanged; recovery attempts still run on every heartbeat tick regardless of that backoff.
- Active work with recent progress logs as