openclaw/docs/tools/loop-detection.md
Ayaan Zaidi 900971b43c
fix(computer): unchanged window reads never reach critical loop detection (#160019)
Closes #159993. Thanks @grape404 for the report.

## What Problem This Solves

Fixes repeated `computer` `get_window_state` calls remaining at loop warnings when an unchanged window returns fresh observation and element references.

## User Impact

With loop detection enabled, unchanged window reads reach the existing critical no-progress threshold. Defaults, thresholds, delivered tool results, fresh references, and stale-reference rejection are unchanged.

## Why This Change Was Made

The computer tool supplies a private stable outcome identity, following the existing progress-card and Code Mode outcome mechanism. Comparison ignores only the observation ID and element references; image data, labels, values, bounds, and other observation metadata remain significant.

No overlap with Pash/Sarah changes.

## Evidence

- Real isolated Gateway agent runs against a scripted OpenAI-compatible model and a synthetic desktop connected over the real paired-node protocol (no real apps controlled): base completed 35 identical reads with `generic_repeat` warnings; candidate blocked the next call after 20 identical outcomes.
- A changing-window control completed 35 reads without a critical block.
- The first 20 model-facing tool result strings matched base byte-for-byte and carried fresh references. After two reads, using the first read's references returned `COMPUTER_STALE_OBSERVATION`.
- Focused observation suite: 14 tests passed. Disabling the owner fix made the unchanged-window regression fail with warning instead of critical; the changing-window control still passed. Initial suite wall time: 24.65 seconds including worker preparation; added cases: 35 ms and 4 ms.
- Native desktop driver behavior was not exercised; the fixture supplied accessibility observations through the actual computer tool and loop detector.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-28 07:19:06 +05:30

8.7 KiB
Raw Blame History

summary title read_when
How to enable guardrails that detect repetitive tool-call loops Tool-loop detection
A user reports agents getting stuck repeating tool calls
You need to control repetitive-call protection
You are editing agent tool/runtime policies
You hit `compaction_loop_persisted` aborts after a context-overflow retry

OpenClaw has two cooperating guardrails against repetitive tool-call patterns, both configured under tools.loopDetection:

  1. Loop detection (enabled) - disabled by default. Watches the rolling tool-call history for repeated patterns and unknown-tool retries.
  2. Post-compaction guard - enabled whenever enabled is not explicitly false. Arms after every compaction-retry and aborts the run if the agent repeats the same (tool, args, result) triple within the window.

Set tools.loopDetection.enabled: false to silence both guardrails.

Why this exists

  • Detect repetitive sequences that make no progress.
  • Detect high-frequency no-result loops (same tool, same inputs, repeated errors).
  • Detect specific repeated-call patterns for known polling tools.
  • Break context-overflow -> compaction -> same-loop cycles instead of letting them run indefinitely.

Configuration block

Global setting:

{
  tools: {
    loopDetection: {
      enabled: false, // master switch for the rolling-history detectors
    },
  },
}

Per-agent override (optional, at agents.entries.*.tools.loopDetection):

{
  agents: {
    entries: {
      "safe-runner": {
        default: true,
        tools: {
          loopDetection: {
            enabled: true,
          },
        },
      },
    },
  },
}

The per-agent setting overrides the global setting.

You can also enable the global rolling-history detectors in Settings → Agent Defaults → Tools in the Control UI. Reset the setting to its default to disable the rolling detectors while keeping the post-compaction guard; explicitly turning it off disables both.

Field behavior

Field Default Effect
enabled false Master switch for the rolling-history detectors. false also disables the post-compaction guard.

Execution titles do not distinguish otherwise identical exec calls. Code Mode also supplies private outcome identities: bookkeeping counters and continuation IDs do not count as progress. Automatically retained values use their original value identity rather than a fresh result-reference ID. Pending work is compared by its operation and arguments, while changed output, returned values, and errors remain meaningful. The displayed receipts and guest data are unchanged, including guest fields named telemetry or pendingToolCalls. This does not detect arbitrary semantically equivalent JavaScript rewrites or make background processes survive a Gateway restart.

For exec, no-progress hashing compares stable command outcomes (status, exit code, timed-out flag, output) and ignores volatile runtime metadata such as duration, PID, session ID, and working directory. For typed terminal failures, it also ignores diagnostic timestamps, explicit attempt or retry counters, elapsed durations, and labeled process IDs. Other text and numbers remain significant, so a new failure cause resets the streak. Outbound message-send results are hashed with volatile per-call ids (message id, file id, timestamp) stripped, so delivery IDs alone do not make repeated equivalent sends look like progress. When a run id is available, history is evaluated only within that run, so scheduled heartbeat cycles and fresh runs do not inherit stale loop counts from earlier runs.

Successful progress_card calls are compared using the saved Markdown and plan, not their write revision or receipt wording. Saved revisions and delivered receipts are unchanged, so a requested refresh still receives a newer saved revision even when the card content is unchanged. Errors and results without the tool’s private semantic outcome keep full outcome comparison.

Window observations from computer get_window_state are compared without fresh observation and element references. Pixels, element labels, values, bounds, and other observation data still count as changes. Model-facing results retain fresh references, and stale references remain invalid for subsequent input.

Outcome comparisons also ignore fresh external-content wrapper nonces, including wrapped errors and JSON results. Delivered security markers remain unchanged; payload text, status, timestamps, and durations still distinguish network results. This is a syntactic comparison: literal or copied text matching the complete wrapper format also ignores nonce-only changes. It does not authenticate content, change authorization, or modify delivered tool results.

  • For smaller models, set enabled: true. Flagship models rarely need rolling-history detection and can leave the master switch unset while still benefiting from the post-compaction guard.
  • To disable everything, including the post-compaction guard, set tools.loopDetection.enabled: false explicitly.

Post-compaction guard

After a compaction-retry following a context-overflow, the runner arms a short-window guard on the next few tool calls. If the agent emits the same (toolName, argsHash, resultHash) triple enough times within that window, the guard concludes compaction did not break the loop and aborts the run with a compaction_loop_persisted error.

The guard is gated by the master tools.loopDetection.enabled flag with one twist: it stays enabled when the flag is unset or true, and only turns off when the flag is explicitly false. This is intentional - the guard exists to escape compaction loops that would otherwise burn unbounded tokens, so a no-config user still gets the protection.

{
  tools: {
    loopDetection: {
      // master switch; set false to disable the guard along with the rolling detectors
      enabled: true,
    },
  },
}
  • The guard compares normalized outcome hashes, not raw result bytes. Meaningful changes keep it from aborting; fresh wrapper nonces alone do not count as progress.
  • It only arms in the immediate aftermath of a compaction-retry, not at other points in a run.
The post-compaction guard runs whenever the master flag is not explicitly `false`, even if you never wrote a `tools.loopDetection` block. To verify, look for `post-compaction guard armed for N attempts` in the gateway log immediately after a compaction event.

Logs and expected behavior

When a loop is detected, OpenClaw logs a loop event and either warns or blocks the next tool-cycle depending on severity, protecting against runaway token spend and lockups while preserving normal tool access.

  • Warnings come first. On OpenClaw-executed tool calls, a short system note is appended to the affected tool result so the model can change approach before a critical block. Warnings share the diagnostic log's rate limit, rather than appearing on every repeated call. The raw outcome is recorded before the note is added, so warning text does not count as progress.
  • Blocking follows once a pattern persists past the warning threshold.
  • Repeating wait with identical arguments and outcomes ten times blocks the next wait. Changed outcomes reset the streak. This uses the same recovery response and terminal handling described below; it does not cancel a tool call that is still executing.
  • In the embedded agent loop, the first critical loop blocks the whole tool batch before any tool in that batch runs. The model then gets one more response with its normal tools.
  • During that response, the model can answer, ask a question, or continue with a different tool or different arguments.
  • Another critical loop in the same run blocks its whole batch and ends the run. A new user run starts with a fresh recovery allowance.
  • The post-compaction guard emits compaction_loop_persisted errors naming the offending tool and identical-call count.
Allow/deny policy for shell execution. Reasoning effort levels and provider-policy interaction. Spawning isolated agents to bound runaway behavior. Full `tools.loopDetection` schema and merging semantics.