openclaw/src/node-host/node-worker-supervisor-contract.ts
Peter Steinberger 8277325ccf
fix(nodes): prevent orphaned work and incompatible launches (#149158)
* fix(nodes): prevent orphaned work after cancellation and crashes

Keep hosted-worker capacity occupied until the process owner confirms descendant cleanup. Preserve completed turn results, peer isolation, and recoverable cleanup ownership when a replacement node host shuts down.

Allow the exact cancelled worker to settle its finishing acknowledgment without recreating publication or execution authority. Carry private lineage descriptors through the existing worker-start channel and package the sealed relay and anchor with the worker bundle.

* fix(nodes): preserve compatible starts across fleet upgrades

Keep released workers on their supported type-only startup and detached process-group owner. Select stronger relay ownership from the worker build capability.

Negotiate captured exec-policy support separately at node inventory so unsupported hosts show the existing update-required outcome before OpenClaw worker dispatch. Preserve remote-exec eligibility, current-authority checks, and cleanup operations.

Validated with released-worker startup and termination proof, registered inventory and dispatch regressions, focused sibling tests, changed checks, and independent review.

* test(nodes): keep current-worker fixtures eligible

Declare captured exec-policy support on the five current-node fixtures that exercise prepared dispatch, replay, authority revocation, and admission. Forward execution mode through the prepared fixture currentness callback independently of slot consumption.

Preserve assertions and deliberate legacy and remote-exec fixtures. All16 reproduced failures are fixed;243 tests across16 files and selected checks pass. Production, build, and dependency inputs are unchanged from97fe.

* ci: refresh node lifecycle merge validation

* fix(node-host): preserve cleanup evidence during worker recovery

Restore released process-group recovery after the leader exits. Record the selected transport before admission and retain exact-anchor root/lineage completion in a journal-owned companion table before releasing capacity after restart.

Use existing-file completion transactions so late cleanup cannot recreate removed state. Preserve public receipts, schema version and terminal retention; drain active modern workers before rollback to an older writer.

* fix(node-host): keep schema DDL annotation beside its call

Preserve the existing canonical first-use schema exception at the leading callsite recognized by the SQL guard. No SQL, guard rule or runtime behavior changes.

* fix(node-host): load journal writer before source helper cleanup

* test(macos): gate readiness recovery after owner failure

Hold the failed startup owner beyond an explicit waiter budget, then require a fresh health request and cleared failure state. Preserve the owner deadline and join cancellation cleanup. Production inputs are unchanged by this commit; disposable native proof runs on a separate task branch.

* fix(nodes): keep free capacity available during recovery

Bound observation of an unfinished cleanup anchor without releasing its reservation or escalating against it. Reconcile retained physical owners through status even when their requested turn has completed, preserving the turn outcome and requiring both lineage completion and tree extinction before freeing capacity.

* fix(nodes): recover capacity after bounded restart observation

Retain exact cleanup observation in the node supervisor after startup returns, publish freed capacity automatically after verified cleanup, and stop and join observation on shutdown. Share cancellation intent and preserve completed turn results.
2026-09-19 05:34:17 -07:00

67 lines
2.4 KiB
TypeScript

import {
parseNodeWorkerSupervisorReceipt,
type NodeWorkerEnvironmentStopInput,
type NodeWorkerLaunchInput,
type NodeWorkerSupervisorIdentity,
type NodeWorkerSupervisorReceipt,
} from "../worker/node-supervisor-protocol.js";
import type {
NodeWorkerWorkspaceRetainInput,
NodeWorkerWorkspaceRetainResult,
} from "../worker/node-workspace-retain-protocol.js";
import type { WorkerConnectionEndpoint } from "../worker/worker-connection-endpoint.js";
import type { NodeWorkerLaunchReceipt } from "./node-worker-launch-store.js";
export {
parseNodeWorkerCancelInput,
parseNodeWorkerEnvironmentStopInput,
parseNodeWorkerLaunchInput,
parseNodeWorkerLookupInput,
} from "../worker/node-supervisor-protocol.js";
export type {
NodeWorkerLaunchInput,
NodeWorkerSupervisorIdentity,
NodeWorkerSupervisorReceipt,
} from "../worker/node-supervisor-protocol.js";
export type NodeWorkerSupervisorControl = {
launch(
input: NodeWorkerLaunchInput,
connectionEndpoint: WorkerConnectionEndpoint,
signal?: AbortSignal,
): Promise<NodeWorkerLaunchReceipt>;
status(launchId: string): Promise<NodeWorkerLaunchReceipt | undefined>;
retainWorkspaces(
input: NodeWorkerWorkspaceRetainInput,
signal?: AbortSignal,
): Promise<NodeWorkerWorkspaceRetainResult>;
cancel(expected: NodeWorkerSupervisorIdentity): Promise<NodeWorkerLaunchReceipt | undefined>;
stopEnvironment(input: NodeWorkerEnvironmentStopInput): Promise<void>;
};
export function projectNodeWorkerSupervisorReceipt(
receipt: NodeWorkerLaunchReceipt,
): NodeWorkerSupervisorReceipt {
const identity = {
launchId: receipt.launchId,
planHash: receipt.planHash,
environmentId: receipt.environmentId,
sessionId: receipt.sessionId,
ownerEpoch: receipt.ownerEpoch,
placementGeneration: receipt.placementGeneration,
runId: receipt.runId,
};
const projected =
receipt.state === "completed"
? { ...identity, state: receipt.state, resultJson: receipt.resultJson }
: receipt.state === "failed" ||
receipt.state === "interrupted" ||
receipt.state === "cancelled"
? { ...identity, state: receipt.state, errorText: receipt.errorText }
: { ...identity, state: receipt.state };
const parsed = parseNodeWorkerSupervisorReceipt(projected);
if (!parsed) {
throw new Error("node worker supervisor durable receipt is inconsistent");
}
return parsed;
}