openclaw/docs/plugins/codex-harness/app-server.md
Ayaan Zaidi 08498ec40d
fix(cron): stale automatic tool lists block scheduled jobs from tools their owner has (#162432)
Related: #130753, #137832, #147969

## What Problem This Solves

Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.

## User Impact

User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.

Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.

## Why This Change Was Made

**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.

**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.

**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.

Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.

## Evidence

Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.

| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |

In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.

Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.

1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:

| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |

3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.

An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.

Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.

## Bounded cost

No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.

LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 00:31:29 +08:00

16 KiB

summary read_when title sidebarTitle
App-server transport, approval posture, auth order, and environment isolation
You are choosing an app-server transport or approval posture
You need the Codex auth selection order
You are isolating the Codex app-server environment
Codex app-server policy App-server policy

How OpenClaw starts and authenticates the Codex app-server, and what it isolates from the operator environment. Part of the Codex harness guide; Where each section moved lists every section.

App-server policy

By default, the plugin starts OpenClaw's managed Codex binary locally with stdio transport. Set appServer.command only to intentionally run a different executable. Verified setup accepts a native Codex executable or the official @openai/codex npm entrypoint, including its installed symlink or Windows npm launcher. Arbitrary wrapper scripts cannot be verified because their native target is unknown; select the native executable or official npm launcher instead. An app-server proxy also cannot supply verified setup because its local executable only forwards requests to a separate daemon. Codex classifies WebSocket transport as experimental and unsupported; use it only for non-production testing against an app-server already running elsewhere:

{
  plugins: {
    entries: {
      codex: {
        enabled: true,
        config: {
          appServer: {
            transport: "websocket",
            url: "ws://gateway-host:39175",
            authToken: "${CODEX_APP_SERVER_TOKEN}",
          },
        },
      },
    },
  },
}

Ask OpenClaw can verify an already configured model through an explicitly configured WebSocket or Unix socket app-server. The initial Codex setup and sign-in flow still requires local stdio; finish sign-in on the remote host and configure the remote endpoint before using this verification path. Remote verification binds the selected endpoint, connection credentials, and initialized Codex identity. It trusts that configured service; it does not attest the remote executable's bytes. OpenClaw rechecks the connection selection before reuse and compares the initialized identity on a new connection before starting a thread. Endpoint, credential, version, or reported Codex home/platform changes require fresh inference verification. Model, authentication, managed requirements, and tool-policy checks still apply.

WebSocket transport proactively establishes the app-server connection at gateway startup and limits the opening handshake to 10 seconds. An idle connection sends a WebSocket ping every 20 seconds and allows 20 seconds for its matching pong. A healthy app-server message or pong resets the missed-heartbeat count; five consecutive missed pongs close the connection. Transient failures reconnect automatically with bounded, jittered exponential backoff. Authentication failures and unsupported app-server versions stop reconnecting and report that operator action is required. Ping and pong frames are transport-level health checks: they do not start a Codex turn or invoke a model. Local stdio and Unix transports do not perform these remote connection checks.

When a caller needs a connection during remote replacement, acquisition makes up to three connection attempts within the caller's timeout and cancellation scope. This applies only when the WebSocket never opened, so no buffered initialization frame reached the server. Authentication and certificate errors fail immediately. Requests on an opened connection, including model turns and tool execution, are not replayed by this recovery.

WebSocket and Unix socket shutdown settles when the connection closes, including when the server disconnected first. If the peer cannot complete the closing handshake, OpenClaw terminates its socket at the shutdown deadline. A closed connection does not prove that work on the remote app-server has stopped.

Local stdio app-server sessions default to the trusted local operator posture: approvalPolicy: "never", approvalsReviewer: "user", and sandbox: "danger-full-access". If local Codex requirements disallow that implicit YOLO posture, OpenClaw selects allowed guardian permissions instead. When an OpenClaw sandbox is active for the session, OpenClaw disables Codex native Code Mode, user MCP servers, and app-backed plugin execution for that turn instead of relying on Codex host-side sandboxing. Shell access instead goes through OpenClaw sandbox-backed dynamic tools such as sandbox_exec and sandbox_process when the normal exec/process tools are available.

Use normalized OpenClaw exec mode for Codex native auto-review before sandbox escapes or extra permissions:

{
  tools: {
    exec: {
      mode: "auto",
    },
  },
  plugins: {
    entries: {
      codex: {
        enabled: true,
      },
    },
  },
}

For Codex app-server sessions, tools.exec.mode: "auto" maps to Codex Guardian-reviewed approvals: usually approvalPolicy: "on-request", approvalsReviewer: "auto_review", and sandbox: "workspace-write" when local requirements allow those values. In tools.exec.mode: "auto", OpenClaw does not preserve legacy unsafe Codex approvalPolicy: "never" or sandbox: "danger-full-access" overrides; use tools.exec.mode: "full" for an intentional no-approval Codex posture. The legacy plugins.entries.codex.config.appServer.mode: "guardian" preset still works, but tools.exec.mode: "auto" is the normalized OpenClaw surface.

For the mode-level comparison with host exec approvals and ACPX permissions, see Permission modes. For every app-server field, auth order, environment isolation, and timeout behavior, see Codex harness reference.

Native approval audit evidence

With tools.exec.mode: "ask" and the Codex user reviewer, native command and file prompts use OpenClaw's two-phase operator approval route. The prompt shows only decisions that the native request can preserve. A command with only a one-shot native decision offers allow-once and deny; byte-bound script approvals also remain one-shot. File prompts support both one-shot and session approval. For commands, Allow Always uses session trust when Codex offers it. Otherwise, it can request a persistent native allow rule when the prompt can show the exact command prefix or network host and its scope across future sessions. Codex owns applying and saving that rule; the approval event reports the requested amendment, not confirmation that it was saved. Automatic command and file approvals remain one-shot and never select a persistent policy amendment.

If another connected Codex client answers a native approval request, OpenClaw dismisses the matching pending prompt without sending a second answer or treating that resolution as a timeout or tool failure.

Terminal operator decisions reuse the Gateway's authoritative approval row and its exact execution binding. When execution identity collection is enabled, inspect the admitted run with openclaw audit --run <run-id> --explain. The resulting receipt can report allow-once, allow-always, denial, no-route, expiry, or cancellation without exposing command text, patch content, paths, or native request ids.

Codex auto-review, full-access policy, and native hook or OpenClaw policy decisions do not create an operator approval row. Missing or stale native turn context is rejected before routing. These cases therefore do not produce an enforced operator-approval receipt; audit inspection does not reconstruct one from later tool events.

Auth order

In the default per-agent home, auth is selected in this order:

  1. Ordered OpenAI auth profiles for the agent, preferably under auth.order.openai. Run openclaw doctor --fix to migrate older legacy Codex auth profile ids and legacy Codex auth order.
  2. The app-server's existing account in that agent's Codex home.
  3. For local stdio app-server launches only, CODEX_API_KEY, then OPENAI_API_KEY, when no app-server account is present and OpenAI auth is still required.

When OpenClaw sees a ChatGPT subscription-style Codex auth profile, it removes CODEX_API_KEY and OPENAI_API_KEY from the spawned Codex child process. That keeps Gateway-level API keys available for embeddings or direct OpenAI models without making native Codex app-server turns bill through the API by accident. Explicit Codex API-key profiles and local stdio env-key fallback use app-server login instead of inherited child-process env. WebSocket app-server connections do not receive Gateway env API-key fallback; use an explicit auth profile or the remote app-server's own account.

If a subscription profile hits a Codex usage limit, OpenClaw records the reset time when Codex reports one and tries the next ordered auth profile for the same Codex run. When the reset time passes, the subscription profile becomes eligible again without changing the selected openai/gpt-* model or Codex runtime.

An exhausted usage percentage does not put a model on cooldown when Codex reports ordinary usage as available or unknown. A known exhausted quota can still supply a scheduled reset hint; that hint does not guarantee renewed availability. Feature-specific resets remain separate from ordinary account permission.

When native Codex plugins are configured, OpenClaw reads and caches one runtime-and-workspace-scoped plugin/installed snapshot. That one snapshot covers configured plugins from Codex-discovered marketplaces, including disabled plugin ownership. plugin/read resolves only explicitly configured plugin details. /codex plugins available queries plugin/list with the bound workspace, while /codex plugins install <plugin>@<marketplace> is the owner- or administrator-authorized installation path. Routine thread setup retains existing explicitly configured curated-plugin recovery.

app/installed supplies the installed app runtime snapshot, and app/read supplies authenticated app metadata in batches of at most 100 app IDs. OpenClaw force-refreshes a cold snapshot once and consolidates successful curated installations into one app-inventory refresh. Ordinary cached reads do not force a connector refresh for every thread.

An authorized app can initially appear disabled or non-callable because Codex has not yet applied the target thread's restrictive app configuration. OpenClaw provisionally admits only explicitly allowed, ownership-proven apps, starts the thread with _default.enabled = false, and reads app/installed once with that thread's ID and forceRefresh: false. Missing, disabled, or non-callable apps produce one warning without blocking unrelated chat or heartbeat runs. Codex still enforces app/tool permissions, managed restrictions, and workspace policy; continuing the conversation does not enable an unavailable app.

The check runs before OpenClaw starts a turn or commits a thread binding. If the snapshot request fails, a persistent provisional thread is deleted and an ephemeral thread is unsubscribed. If cleanup cannot be confirmed, OpenClaw retires the app-server connection instead of reusing an unsafe thread.

Account-wide app access never overrides an explicitly disabled configured workspace plugin. When app/read omits that plugin's ownership, OpenClaw uses the plugin/installed snapshot and reads only the exact configured plugin's details to keep its apps denied. This check never installs, enables, or authenticates the plugin.

OpenClaw does not install unknown apps or let the model authorize new plugin installs. Owner-approved plugin installation refreshes the target runtime inventory. Missing inventory methods, authentication errors, transport failures, and connector refresh failures fail closed.

Scheduled app authority

When a Codex creator turn captures scheduled app authority, an automation without an explicit toolsAllow list saves that turn's callable tools and app policy. With a prepared ChatGPT profile, scheduled app access remains bound to that exact profile and account. Without a prepared profile, an agent-scoped configured WebSocket app-server owns the schedule through its connection fingerprint. Reauthenticating that same endpoint to another account does not revoke the schedule: subsequent runs use the endpoint's current account, subject to the captured app ceiling and current app/tool policy. Scheduled authority does not store or replay authentication credentials.

Scheduled app approval ceilings preserve native tool overrides and the approval policy of the account identified by each tool. For tools that select an account when called, the shared tool ceiling uses the strictest combination of the configured account and default policies. Such tools can require approval across accounts even when one account permits the action automatically.

Removing or un-configuring the endpoint, changing its connection fingerprint, or changing its captured managed requirements rejects the run before app execution. The job remains inspectable, with an error in automation run history and its last-run state; normal failure backoff still applies. Restore the authorized connection or recreate the automation from a fresh authenticated owner turn. Account changes that remove access to a captured app also fail visibly.

Before rolling back to a build without configured-endpoint authority and cron authority hydration, disable these jobs with openclaw automations disable <id> and verify them with openclaw automations list --all. Do not rely on an older binary to enforce the new authority envelope. Keep the jobs disabled until you return to a supporting build or recreate them under that build's supported auth path. See Automations for run history and failure handling.

Environment isolation

For local stdio app-server launches, OpenClaw sets CODEX_HOME to a per-agent directory so Codex config, auth/account files, plugin cache/data, and native thread state do not read or write the operator's personal ~/.codex by default. OpenClaw preserves the normal process HOME; Codex-run subprocesses can still find user-home config and tokens, and Codex may discover shared $HOME/.agents/skills and $HOME/.agents/plugins/marketplace.json entries. With appServer.homeScope: "user", OpenClaw instead uses the native user Codex home and its existing account without injecting an OpenClaw auth profile. Canonical openai/* chats on a user-home stdio or Unix connection also retain the native configured model provider; select the model with the canonical OpenClaw model ref. Explicit non-OpenAI providers remain explicit. Prepared route compatibility and subscription/API-key account checks still apply.

If a deployment needs additional environment isolation, add those variables to appServer.clearEnv:

{
  plugins: {
    entries: {
      codex: {
        enabled: true,
        config: {
          appServer: {
            clearEnv: ["CODEX_API_KEY", "OPENAI_API_KEY"],
          },
        },
      },
    },
  },
}

appServer.clearEnv only affects the spawned Codex app-server child process. OpenClaw removes CODEX_HOME and HOME from this list during local launch normalization: CODEX_HOME stays pointed at the selected agent or user scope, and HOME stays inherited so subprocesses can use normal user-home state.

Verified local setup turns also attest the selected Codex launcher and package. Inherited NODE_OPTIONS may contain bounded resource, warning, DNS result order, network-family autoselection, environment-proxy, and CA-source options because those settings cannot preload code or change module resolution. For example, --dns-result-order=ipv4first --no-network-family-autoselection is allowed. Malformed or unknown options and code-loading options such as --require or --import fail closed. If an inherited option is not needed by Codex, remove NODE_OPTIONS with appServer.clearEnv.

Local testing env overrides

  • OPENCLAW_CODEX_APP_SERVER_BIN bypasses the managed binary when appServer.command is unset.
  • OPENCLAW_CODEX_APP_SERVER_ARGS accepts a quoted argument string; see argument parsing.
  • OPENCLAW_CODEX_APP_SERVER_MODE=yolo|guardian
  • OPENCLAW_CODEX_APP_SERVER_APPROVAL_POLICY
  • OPENCLAW_CODEX_APP_SERVER_SANDBOX

OPENCLAW_CODEX_APP_SERVER_GUARDIAN=1 was removed in 2026.4.22. Use plugins.entries.codex.config.appServer.mode: "guardian" instead, or OPENCLAW_CODEX_APP_SERVER_MODE=guardian for one-off local testing. Config is preferred for repeatable deployments because it keeps the plugin behavior in the same reviewed file as the rest of the Codex harness setup.