* feat: show rolling session recaps in Activity * fix: respect lifecycle and viewer permissions for recaps
28 KiB
| summary | read_when | title | ||
|---|---|---|---|---|
| Which SQLite database holds what, and the tables behind individual features |
|
Database layout |
Database layout
| Scope | Default path | Contents |
|---|---|---|
| Global control plane | ~/.openclaw/state/openclaw.sqlite |
Shared configuration state, registries, approvals, plugin state, and shared runtime state |
| Per-agent data plane | ~/.openclaw/agents/<agentId>/agent/openclaw-agent.sqlite |
Sessions, transcripts, memory indexes, auth state, conversation state, and agent-scoped runtime state |
The task registry uses the shared state database. Runtime trajectory events live with their sessions in the per-agent database or a configured shared session SQLite store.
Activity session recaps
Activity stores one optional activitySummary object in the existing session_nodes.entry_json session metadata. This is a reconstructible cache; the transcript remains canonical. The approved persistence design adds no SQL table, column, or database schema-version change. Current and v2026.9.4 metadata serializers preserve unknown optional fields; unknown recap payload versions are treated as cache misses.
Payload version 1 records the recap text, generation time, session ID and lifecycle revision, transcript generation and leaf, chronological coverage, and whether oversized message content was omitted. A rewind or replacement invalidates an incompatible source binding. The Gateway reads bounded transcript chunks outside the metadata write and rechecks the current lifecycle and transcript branch before committing. Recap writes preserve session activity timestamps and ordering.
The latest recap survives restart and archival. Deleting the session removes it; reset or replacement makes the prior lifecycle's recap unusable. Incognito sessions do not persist or generate this cache. A shared, bounded Gateway queue deduplicates generation across viewers, retains the previous recap on failure, and uses only the configured utility route. Disabling that route stops new generation. Removing or ignoring the optional field is a rollback path that leaves session and transcript data intact; removing the feature does not require reversing a database migration.
Cold transcript archives
The per-agent session_transcript_cold_archives table records cold transcript
locations alongside session_windows and transcript_events. Each row belongs
to a retained session window and identifies its generation, archive name, hash,
counts, and sizes. The payload lives in an immutable compressed JSONL file, or
in the row's blob when embedded by a supported backup.
The default archive directory is
~/.openclaw/agents/<agentId>/sessions/cold/, with filenames
<sha256>.jsonl.zst. A database in a directory named agent uses its sibling
sessions/cold/ directory; other store layouts use cold/ beside the database.
These files contain authoritative history. See
cold transcript storage
for retention and restoration, and
agent schema 20
for the schema and update contract.
Plugin state listing index
Plugin keyed stores use the shared plugin_state_entries table. Its listing
index includes expires_at after the existing plugin, namespace, creation-time,
and entry-key columns, so live-row counts can read the index without fetching
stored values. Quotas, TTL cutoffs, ordering, and row contents are unchanged.
Writable startup and openclaw doctor --fix replace the older four-column
definition through canonical index repair, without a schema-version bump. The
repair builds temporary indexes and runs the existing table and full-file
integrity checks; allow for extra disk space and work proportional to stored
entries during the first repair.
An older build can rebuild the same index back to its expected definition. Full-schema read-only validation rejects a mismatched definition until a writable owner repairs it; lightweight readers that validate only the numeric schema version may read either shape. See the accepted index design for upgrade, reverse-repair, and performance proof requirements.
Mentions Inbox
The mentions Inbox uses existing
config_machine_state rows in state/openclaw.sqlite.
notifications.mentions.source.* records retain typed source identities,
recipients, mention identifiers, expiry times, and dismissal bookkeeping;
notifications.mentions.head records the revision and sequence. Writes use the
existing table and primary key, with no new tables, columns, indexes, or schema
version change.
Retention remains seven days from creation, capped at 100 entries per profile, 10,000 entries globally, and 10,000 source identities for duplicate suppression. Restarts preserve retained entries, dismissals, and their original expiry times. Loading stored state does not replay browser notifications or scan transcripts to reconstruct old mentions.
ACP replay accounting
The shared acp_replay_sessions and acp_replay_events tables retain bridge
replay history. Their estimated_bytes columns count the UTF-8 bytes of each
persisted text field, plus 32 bytes per row. Session totals include their events.
This is a retained-content estimate, not a limit on SQLite file, page, or WAL size.
Older releases counted characters inconsistently, undercounting Unicode and allowing unchanged metadata writes to drift. The existing app-version upgrade repair and explicit shared-state schema repair rebuild all derived totals atomically, preserving event JSON text, identifiers, timestamps, and sequence. Repair does not prune history. The next ordinary session write applies the existing caps and eviction order, so corrected Unicode history may trim sooner and use transcript fallback when loaded.
A current-app-version reopen skips this repair. Replacing code without changing the app version does not repair an already-open or current-version database; explicit schema repair remains the repair owner for that case. Accounting repair cannot recover history already evicted by an older writer. See ACP CLI.
Meeting transcript tables
Meeting captures use three STRICT tables in the shared
state/openclaw.sqlite database, separate from per-agent conversation transcripts.
The transcript store (src/transcripts/store.ts) owns their reads and writes;
src/transcripts/sqlite-schema.ts ensures the tables on first use. Markdown and
JSON files under the transcripts directory are explicit exports, not runtime
storage. See Transcripts CLI.
meeting_transcript_sessions
One row per capture identity. The primary key is (session_id, started_at);
selector is unique. Indexes support start-time, session-ID, slug, and export-key
lookups.
| Columns | Type | Purpose |
|---|---|---|
session_id, started_at |
TEXT NOT NULL |
Capture ID and original start time. |
selector, export_key, session_slug |
TEXT NOT NULL |
Canonical selector and derived export identity. |
provider_id, source_json |
TEXT NOT NULL |
Source provider and locator. |
title, stopped_at, metadata_json |
Nullable TEXT |
Display title, terminal time, and session metadata including ownership. |
export_manifest_json |
TEXT NOT NULL, default {} |
Export artifact ownership manifest. |
export_pending_json |
TEXT NOT NULL, default [] |
Pending export artifacts. |
next_utterance_seq |
Nonnegative INTEGER NOT NULL, default 0 |
Next append sequence. |
created_at_ms, updated_at_ms |
Nonnegative INTEGER NOT NULL |
Store timestamps. |
Reopening an occupancy-driven capture clears stopped_at without changing the
primary key, so the same meeting retains its utterances.
New transcript admissions record sessionIdOrigin (generated or supplied)
in metadata_json. The store preserves that value, including its absence or
invalidity in legacy rows, on later writes to the same primary key. Occupancy
reopening requires an explicitly generated origin; an unknown origin starts a
fresh capture and leaves the old record intact. The existing newest-candidate
query and ten-minute window are unchanged.
This adds no schema, index, version, or backfill. Doctor metadata restoration preserves an explicitly recorded origin and leaves unknown origins unknown. Older runtimes do not enforce this rule, so downgrading also removes the fixed-ID history protection. See the accepted ID-origin decision.
meeting_transcript_utterances
Append-ordered speech records. The primary key is
(session_id, session_started_at, sequence); the session pair references
meeting_transcript_sessions(session_id, started_at) with ON DELETE CASCADE.
| Columns | Type | Purpose |
|---|---|---|
session_id, session_started_at |
TEXT NOT NULL |
Owning capture identity. |
sequence |
Nonnegative INTEGER NOT NULL |
Stable append order within the capture. |
utterance_id, started_at, ended_at |
Nullable TEXT |
Provider utterance identity and timing. |
speaker_id, speaker_label |
Nullable TEXT |
Provider speaker identity and display label. |
text |
TEXT NOT NULL |
Captured transcript text. |
final |
Nullable INTEGER, 0 or 1 |
Whether the provider marked the utterance final. |
metadata_json |
Nullable TEXT |
Provider utterance metadata. |
meeting_transcript_summaries
One current summary per capture. The primary key is
(session_id, session_started_at) and references the session primary key with
ON DELETE CASCADE. At least one of summary_json or markdown must be non-null.
| Columns | Type | Purpose |
|---|---|---|
session_id, session_started_at |
TEXT NOT NULL |
Owning capture identity. |
generated_at |
Nullable TEXT |
Summary generation time. |
summary_json |
Nullable TEXT |
Free-form summary, including participants, source (model or heuristic), and optional model reference. |
markdown |
Nullable TEXT |
Rendered meeting notes. |
utterance_count |
Nonnegative INTEGER NOT NULL |
Number of utterances covered by the stored summary. |
These are existing feature-local tables. Occupancy episodes and model-backed notes do not change their schema or database version.
Update run ledger
update_runs stores one durable record per update in the shared
state/openclaw.sqlite database. src/infra/update-run-ledger.ts owns writes
from the admitting Gateway, orchestrator CLI, and restarted Gateway. The table
is additive at shared schema version 15: the canonical schema declares it and
first use ensures it inside the same write transaction. Existing tables and the
schema version stay unchanged; older readers ignore the new table.
run_id is the UUID primary key. Rows retain creation/update timestamps,
trigger, phase, status, reason, origin, target, before/after versions, steps,
verification facts, repair attempts, confirmation/finish timestamps, and known
downtime. Each JSON column has a 16 KiB hard limit with deterministic truncation
and redaction. The ledger stores bounded diagnostic summaries, not raw logs or
credentials. There is no automatic history deletion.
New drivers store optional origin.driver fields host (the hostname), pid,
and startIdentity (the operating system's process-start identity as a decimal
string) in the existing origin_json column. Each adopter becomes the current
driver and retains distinct earlier identities in origin.previousDrivers.
There are at most eight identities in total. Only positively dead identities
are pruned; adoption is refused rather than dropping a live or uninspectable
driver at capacity. If local process identity cannot be captured, adoption
continues with one warning and a retained driver:identity-unavailable step.
That marker permanently excludes the run from automatic reconciliation, even
if known parents later exit; existing recorded identities remain protected.
A fresh run without identity follows the legacy explicit repair/supersession
rules below.
This is additive JSON metadata;
there are no new columns, tables, or schema versions. The separate
verification.pid still identifies the Gateway service, not the updater.
Adoption records a retained driver:adopted step. Detached children can outlive
their parent, so either lifetime can prevent reconciliation. Adopting a terminal
run is refused. Long command and finalization phases renew updated_at_ms
every 30 seconds; only current or retained identities may renew a row. Heartbeat
write failures warn once per driver run and do not abort commands or finalization;
step and outcome writes retain their existing failure behavior. Encoding
reserves space for exact identity bytes before bounding and redacting other
origin diagnostics.
The ledger owns abandonment classification and terminalization. Automatic
recovery requires more than 30 minutes since both updated_at_ms and the latest
step timestamp, plus positive evidence that every recorded driver is dead on the
same host: its PID is gone or its process-start identity differs. Unreadable and
foreign-host identities are inconclusive. The Gateway performs reconciliation
at startup and on active-run polls, rechecking the current row and process
identity in the terminal write transaction. The shared 30-minute constant also
owns the older-updater schema-publication bound described under
Schema bumps and older updaters.
Reconciliation writes status failed, reason abandoned, and a retained
reconcile:abandoned step whose detail names inactive-driver-dead or
operator-reconciled-inactive-run. All unfinished steps become terminal, and
history is retained. Explicit update repair can reconcile inactive identityless
rows when the current Gateway generation is healthy and no post-core repair is
pending. When every recorded driver is positively dead and no
driver:identity-unavailable marker exists, explicit recovery does not require
the inactivity window. It cannot override a live or inconclusive recorded driver. The
2026.9.2 updater
does not record adoption: package-manager and registry preflight can
leave a live updater at its single requested/in_progress step. Older writers
may drop unknown driver JSON fields; identityless rows normally require explicit recovery.
An untouched legacy admission expires automatically after more than 24 hours:
it is still requested / running, has identical creation and update timestamps,
no finish timestamp, no recorded driver, and only its initial requested step.
The ledger retains it as failed with reason legacy-driver-expired and a
reconcile:abandoned step. Startup, update status, status, and the Control UI's
update reads reconcile this shape through the same transaction. Status and failure
reports explain that the update never progressed and recommend openclaw update
to retry. Status retains the latest such advisory even after a newer update finishes.
This fixed legacy expiry does not establish process death. It is the bounded
recovery policy for 2026.9.2-era orphan admissions. Younger rows, progressed rows,
recorded drivers, and retained recovery descriptors keep their existing protections.
Recording driver:identity-unavailable is itself a mutation, so that adoption
cannot match an untouched admission. All other update status history remains
read-only; the existing 30-minute inactivity window is unchanged.
Gateway update reads also run the existing automatic dead-driver reconciliation,
so Control UI admission uses the same recovery decision as its startup watcher.
Explicit new CLI update admission can supersede a legacy row only when it is
the sole running row, has no current or previous driver identity, and exceeds
the same inactivity bound. The transaction finishes it as failed with reason
superseded and a retained reconcile:superseded step whose detail is
operator-started-update-supersedes-inactive-identityless-run, then creates the
new row. This includes dry-run admission, but excludes inherited continuations
and campaigns. abandoned and superseded are additive values in the existing
free-text reason contract. Neither recovery path deletes history.
Successful ledger-only repair records a retained reconcile:acknowledged step.
A terminal abandoned row can substitute for full repair only once, within
30 minutes of its finish time; later repair invocations keep normal plugin
convergence behavior.
Repair also inspects newer failed/abandoned history for unacknowledged post-core
work, regardless of its age. An older active row cannot hide that work. If the
bounded history prefix does not reach the selected recovery rows, repair uses
full finalization rather than claiming that no post-core work remains.
When full finalization is required, the selected inactive rows are rechecked
and reconciled only after successful convergence, before success output.
Explicit recovery validates and commits its selected rows in one transaction;
renewed activity in any selected run preserves the entire selection. Ledger-only
repair also refuses the write if another active run falls outside that selection.
Finalization (finalize:*) and post-update verification markers survive step-count
and diagnostic-byte eviction because repair relies on that history. If retained
metadata alone exceeds a hard limit, the write fails without changing the row.
The CLI and Gateway share WAL-backed transactions, including while the Gateway
is stopped. The first terminal outcome wins; subsequent verification can enrich
its observed facts without rewriting success, failure, skip, or rollback status.
The restart sentinel carries stats.runId and remains the continuation owner;
consuming it does not delete the run row. Chat, CLI, and status reports read that
row. See Run history and reports.
Update installation control
Managed update leases use a separate machine-local managed-update-handoffs.sqlite
database under the secure OpenClaw temporary directory. This owner must remain
available while an update replaces an installation or changes its runtime state.
The update history above remains in the profile's shared state database.
First creation exclusively creates a private file with one filesystem link. Concurrent initializers use that same file without replacing it. SQLite commits the existing schema atomically; ordinary lease inspection reports no lease while the first-use schema is empty. Reads of an absent database create no state. Existing lease rows, claim transactions, and the rule that one updater owns an installation are unchanged.
The normal handoff parent prepares this database before launching its sealed helper. The helper receives the captured database identity and operates only on that existing database, without resolving installation packages or recreating missing or empty state.
File creation applies private permissions before SQLite opens the file, including a protected ACL on Windows. Initialization follows the existing directory-durability policy and does not require the optional fs-safe native binding. Failure stops lease admission before its operation runs. After an interrupted first creation, the normal owner can finish initialization through its existing empty-database recovery path; committed rows remain governed by SQLite's normal transactions. This change requires no schema migration. See the accepted initialization design.
Managed worktree acceleration templates
Managed worktree acceleration uses the first-use worktree_templates table in the shared state database. Each row records a reconstructible source template: repository and Git common directory, destination root, filesystem backend, artifact path, source commit, checkout content key, preparation status, and creation and last-use timestamps. The cache key allows one template per repository and destination root. The template contains no provisioned ignored files or repository setup output.
The worktree service owns template creation, reuse, invalidation, and cleanup under its existing allocation lease. It reserves a preparing row before creating the artifact and publishes ready only after preparation completes. Durable mutations recheck the lease inside synchronous state transactions; filesystem work runs outside those transactions. Cleanup uses the reserved template ID so an old operation cannot delete its replacement. Templates are replaced when the commit or checkout policy changes and retired after seven days without use.
The additive table is ensured on first use and does not change the numeric database schema version. Existing worktree and snapshot records retain their meaning; no existing checkout is migrated or moved. Template artifacts are reconstructible, while registered worktree contents and recovery snapshots retain their existing preservation rules.
Cloud repository workspaces
Repository-only cloud sessions use the first-use session_repository_workspaces table in the shared state database. The existing session entry carries only repositoryWorkspaceId; the shared row owns the canonical agent/session key, repository URL, requested ref, session branch, setup intent, pinned base commit and manifest, accepted checkpoint pointer, and revision. Session reset preserves this owner; a fork receives a distinct owner.
github_repository_publication_requests records shared and personal publication against an immutable accepted checkpoint and the session's admitted lifecycle revision. Reset preserves the session ID and repository checkpoint but invalidates publication authorized before that reset. Personal requests also retain the selected profile and connection generation and require same-owner confirmation after an interrupted publication. Pending publication keeps its original source even after an explicit move materializes a Gateway worktree.
Both tables are additive, lazily ensured on first use, and leave the numeric database schema version unchanged. That is not a compatibility promise for older cloud-session implementations: run a build that understands repository-only sessions when using this state. Existing local managed-worktree sessions keep their existing representation.
Checkpoint Git artifacts live under state/repository-workspaces/<workspace-id>.git, next to the shared database. These are bare repositories containing complete file manifests, cumulative changed-file blobs, and publication snapshots; they are not working checkouts or a backup of upstream Git history. Restoring an entire checkout still requires access to the pinned upstream commit. Back up these artifacts together with the shared and per-agent databases.
Accepted checkpoint history and publication source artifacts remain until explicit session deletion, including after Stop, archive, reset, or Gateway restart. There is no timed checkpoint expiry. Deletion retires publication requests and source ownership before removing their artifact repository; failed cleanup is reported. The managed-worktree idle cleanup and snapshot retention rules do not apply to these checkpoints.
Sandbox runtime reservations
The existing sandbox_registry_entries table owns runtime identity and cleanup.
Backends that opt into reservation persist a generation before provider allocation.
The entry payload retains the original workspaceDir and records runtimeState
as pending, ready, removing, or removing-pending;
no schema version, table, or column is added. Older entries are adopted on first
use, and backends without reservation keep their existing registry behavior.
Factories receive the reserved workspace on replay and shared-scope reuse, rather
than the latest caller's local workspace. Provider execution and repository-scoped
cleanup therefore use the same original owner.
Reservation and publication use synchronous SQLite transactions. Provider work runs outside the transaction under a per-runtime file lock beside the shared database. Lock contenders wait up to 15 minutes, covering the backend's warmup and inspection budgets. Concurrent creators reuse the same generation. Failed provisioning retains its pending ID for replay after restart. Recreate and prune record removal intent before waiting for provisioning, then remove the provider runtime before deleting the row. Cleanup failures retain that intent for retry; stale handles cannot publish readiness or start new operations after removal begins.
removing-pending preserves the fact that provisioning never published readiness.
If ordinary cleanup fails, the Crabbox adapter can replay that same fixed ID from
its original workspace and then release it. This also covers a failure before
Crabbox recorded the request: recovery may allocate and immediately release the
reserved runtime. Unknown outcomes retain the recovery row; an absent local claim
or an error message is not proof that provider resources are absent.
The reservation is canonical recovery state. Do not delete it to clear a provider error. Before downgrading to a version without reservation support, disable the backend and reconcile its pending leases using the current version. Older readers can open the database but do not implement this lifecycle.