Second candidate on the v6.2.0 line, superseding v6.2.0-rc.1 and still
following stable v6.1.2. This is a hardening and bugfix candidate:
auto-update reliability including in-place migration of already-deployed
update units, PBS backup attribution across multiple Proxmox clusters,
Proxmox installer registration including per-canonical-type bootstrap
grants on combined PVE+PBS hosts, Patrol readiness streaming transport
plus a completed-verdict readiness gate, discovery-policy DNS and SSH
retry churn, explicit per-metric threshold off toggles, request-derived
SSO callback URLs, and the new Entra ID SSO guide. Version pins move to
6.2.0-rc.2 across the repo root, Docker bootstrap defaults, and Helm
metadata per the deployment-installability contract; stable install
pointers remain on v6.1.2 until governed promotion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A host running both PVE and PBS is a deployment the docs call officially
supported, and the agent's RunAll registers each product in turn from the
one install token. The bootstrap grant recorded consumption per TOKEN, so
the PVE leg spent it, the PBS leg came back canRegister=false, the agent
wrote a proxmox-pbs-registration-blocked marker, and install.sh printed an
ERROR banner over a PVE source that had registered perfectly well.
- server: consumption is now recorded per canonical type. One PVE create
and one PBS create per token, each still one-shot — a second create of
the same type takes the same 403. The bounds that are not per type stay
singular: the 24h mint-age clock and the first-use bound_hostname are
shared, so whichever type registers first pins the hostname for both and
the second type cannot be aimed at another machine. The per-type ledger
lives in proxmox_registration_consumed_types; a record carrying only
proxmox_registration_completed=true predates it and still reads as every
type consumed, so upgrading cannot revive a token already spent in the
field. The unsuffixed completion block keeps tracking the most recent
completion, which also leaves an older binary reading the same store
failing closed.
- rollback: the consume-before-persist undo is scoped to the keys one
consumption wrote, so a failed PBS source save restores the PBS grant
without resurrecting the PVE grant that already produced a source.
- agent: RunAll no longer lets one product's failure speak for the host. It
attempts and returns the remaining products, errors only when every
detected product failed, and publishes the detected products in a
proxmox-detected-types state marker.
- installer: report_proxmox_registration_outcome reads that marker, waits
for an outcome from each detected product, and prints a success or denial
line per product instead of one verdict. Agents predating the marker keep
the old first-outcome-wins timing so a single-product host does not wait
out the window. The blocked-marker path now only fires on genuine refusals.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The "Patrol tools" readiness check read the cached model-readiness
snapshot's tool-protocol dimension on its own. Since the interrupted-run
handling landed (8d0d74e35, b78330405), a run cancelled after every tool
scenario already passed keeps ToolProtocol at pass while the overall
status reports not_assessed, so the check reported "Patrol ready" from an
evaluation that never completed.
The check now requires the snapshot's own overall verdict (Success)
before reporting ready. A snapshot carrying no verdict at all — overall
status not_assessed, or the interrupted or internal_error cause — is not
turned into a failure either: it falls back to the base-config classifier
exactly as an absent snapshot does, capped at a warning. That cap matters
because not_ready is a blocking status in this payload: it clears
readiness.ready, which disables the Patrol run control in
usePatrolIntelligenceState and drops the page into the setup-only view.
#1640 promises a severed or cancelled check never blames the model and
never blocks Patrol from running in Watch mode, and the runtime gate on
POST /api/ai/patrol/run (PatrolRuntimeReadiness) already treats an
unassessed mode as a warning, so a blocking tools check would have
contradicted the route that actually runs Patrol.
A completed run whose tool protocol passed while the overall verdict fell
short now warns instead of claiming ready. It must not block either: the
dimension that actually failed carries the verdict on its own check
(context quality blocks, latency warns), so blocking here would have
turned today's latency warning into a hard stop.
Regression tests: internal/api/issue1640_readiness_gate_test.go covers
the gate across interrupted, internal-error, completed-pass,
completed-fail, and short-of-pass snapshots, asserting the resulting
runnability of the readiness payload;
internal/ai/issue1640_readiness_gate_test.go drives a real evaluation
that is cancelled at the continuation probe to produce the
ToolProtocol=pass / status=not_assessed snapshot end to end and pins that
PatrolRuntimeReadiness keeps Patrol runnable. The new API test file is
registered in the subsystem verification registry.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two failures on c3fb35c8f. The subsystem lookup unit tests hold their own
hardcoded copies of the settings-shell-and-framing exact_files list, a
third snapshot surface beyond the guard test's, and registering the new
banner proof file left them stale. Synced all eight copies.
The apiClient message-precedence change broke a pre-existing pin that a
short plain-text body outranks the caller fallback on retryable errors
(useReportingPanelState). Body-over-fallback was the long-standing
behavior; what #1640 actually required was dropping markup and oversized
bodies, which stays. Restored body-wins for sane plain text, flipped the
two precedence tests introduced alongside the change, and corrected the
cloud-paid contract paragraph to describe the real order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to 8d0d74e35. The keepalive mechanism was right, the edges
were not.
1. The evaluation ran on a bare goroutine with no recover, so a panic in
provider streaming or validation took the whole Pulse process down.
Before that commit the same panic was on the request goroutine and the
recovery middleware turned it into a logged 500. The goroutine now
recovers, logs the panic with its stack, and answers with an ordinary
readiness result carrying the new internal_error cause and every
dimension reported as not assessed. A Pulse defect is not a model
verdict.
2. Headers were only Set, never committed, despite the comment, the
commit message, and api-contracts.md all claiming otherwise. The
status line went out with the first keepalive at +10s, so a proxy
with a sub-10s time-to-first-byte budget still severed the request.
The transport now writes and flushes WriteHeader(200) before the
ticker starts, matching the pattern the file already uses for SSE.
3. The flusher was resolved with a discarded ok, so a writer that is not
an http.Flusher silently buffered the keepalives and degraded back to
the original bug. It is now checked and logged; the response still
completes, so a warning is the right level here rather than the hard
failure the SSE handlers use.
4. TestIssue1640HandlerUsesKeepaliveTransport grepped the handler source
for substrings, which proves nothing about behaviour. Replaced with a
real httptest.NewServer test that runs a 300ms evaluation and asserts
the client sees the 200 and a body byte before the evaluation
completes, and that the padded body still parses as the expected JSON.
Added coverage for the panic path and the non-flushable writer, and
fixed the eager body[:1] that would panic when a transport regression
left the body empty.
5. The settings readiness banner had no not_assessed branch, so an
interrupted run still rendered the red "Patrol model not verified"
headline: the exact blame-the-model presentation the backend fix
removed. Tone and headline are now exported pure functions with a
neutral treatment for not_assessed and interrupted results, and an
interrupted run cannot claim verification from a max_verified_mode
recorded before the cancellation.
6. createAPIErrorFromResponse let a short plain-text body override an
explicit caller fallbackMessage. A caller passing a fallback knows
which operation it was performing; an intermediary writing the body
does not. Precedence is now canonical JSON, then caller fallback,
then body, with the HTML and oversize suppression unchanged.
7. patrolRunCancelled classified on the raw "context canceled" substring
as its first switch case. Ollama embeds that phrase in its own error
body when it aborts an upstream request, so a genuine provider
failure on a healthy run was classified interrupted and finish()
persisted it as not_assessed. Cancellation is now established from
the run itself (errors.Is(err, context.Canceled), or a cancelled run
context), never from error wording, and the readiness paths classify
through a context-aware entry point. context.DeadlineExceeded keeps
its provider-path timeout classification.
The readiness gate in HandlePatrolModelReadiness keying off ToolProtocol
alone is untouched, as agreed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adversarial review of 9db25ba60 found four residual defects in the
auto-update asset install, plus a doc line it left contradicting itself.
- install_auto_update_assets copied the bundled helper into the staged
mktemp file with an unchecked cp, and both call sites invoke the
function under `if !`, which suppresses errexit for its whole body. A
failed copy (ENOSPC, EIO) fell through to
configure_auto_update_script_repo, whose awk emits a lone GITHUB_REPO=
line for empty input, so a shebang-less one-line stub replaced the
working helper with a "script" that only ever exits 0 - silently
disabling unattended updates. Check the cp, and refuse the swap unless
the staged helper is non-empty and starts with #!.
- Both units were rendered with a bare truncating `cat > "$unit"` whose
status was never checked, and the function's last statement is
safe_systemctl daemon-reload, which returns 0 by design. A failing
write truncated a working unit and still reported success. Render each
unit to ${path}.tmp and commit it with a checked rename, so a failure
leaves the installed unit byte-identical.
- The widened ReadWritePaths could not reach deployed boxes: the unit
that grants the write access is itself the file that has to be
rewritten, and on an existing install the sandbox running the
installer excludes /etc/systemd/system and /usr/local/bin (EROFS). The
Go update pipeline cannot carry it either - pulse.service runs as
User=pulse with its own ProtectSystem=strict over the install and
config dirs only. So probe each destination directory up front and,
when one is blocked, re-exec this already-signature-verified installer
through systemd-run with a new internal --repair-auto-update-units
entry point: PID 1 forks the transient unit, so it starts in the host
mount namespace instead of inheriting the sandbox. The installer is
copied into the install dir first because the calling unit's
PrivateTmp=yes hides its /tmp copy from PID 1. The escape needs root
and systemd-run, and never recurses.
- Keep the ReadWritePaths entries as directory grants: every write now
commits with a rename from a sibling staging file, and rename needs
write access on the containing directory, so the file-level entries
systemd would otherwise accept cannot work. Document the tradeoff in
the unit and the subsystem contract instead.
The deployment-installability contract still claimed the update sandbox
leaves "only the install dir, config dir and /tmp" writable, which the
paragraph the same file gained in 9db25ba60 contradicts; the same stale
rationale had been copied into two test comments.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Settings > Infrastructure installer mints generic host install
tokens, but install.sh auto-detects Proxmox and the agent presents type
pve/pbs at /api/auto-register. The bootstrap grant required an exact
install_type match, so every generic install on a Proxmox node was
denied source creation and the denial was a single buried journal warn.
Four-part fix (#1644):
- server: extend the one-shot bootstrap grant to host-issued install
tokens presenting a canonical Proxmox type. Typed tokens stay pinned,
the grant keeps its settings-write mint requirement, first-hostname
binding, serialized completion, and single consumption across types.
- agent: a canRegister=false denial now logs at error level, returns a
setup error, and records the operator-facing reason in a
proxmox-<type>-registration-blocked state marker.
- installer: report the Proxmox registration outcome in install output
by reading the registered/blocked markers, and poll the server lookup
for a bounded retry window before warning that registration was not
confirmed (readyz flips before the first report cycle).
- setup script: the auto-register transport now captures the HTTP
status alongside the body (no -f), making the invalid-setup-token
branch reachable via 401/403 instead of a dead server-string grep,
and operator guidance names Settings -> Infrastructure instead of the
retired Nodes page (also updated in docs/PBS.md and the pinned
assertions in contract, setup-script, and repoctl docs tests).
Regression proof: internal/api/issue1644_host_install_token_proxmox_test.go
plus new install.sh proofs for the retry window and blocked-marker
surfacing, and the updated hostagent blocked-registration test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The registry gained internal/monitoring/issue1638_dns_cache_test.go
(108aa4e20) and internal/api/issue1640_readiness_transport_test.go
(8d0d74e35) as registered verification files, but the expected
verification-requirement snapshots in canonical_completion_guard_test.py
were not updated alongside them, so every Canonical Governance run on
main has been failing its guard unit-test step since. Add both files to
the expected lists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Residual auto-update defects found while triaging #1643 and #1637
(the primary regression was fixed in 806cbe83d):
- The generated pulse-update.timer carried both OnCalendar=daily and
OnCalendar=02:00, so with RandomizedDelaySec=4h every box attempted
two updates per day (00:00-04:00 and 02:00-06:00 windows). Keep the
single documented 02:00 schedule.
- The generated pulse-update.service sandbox (ProtectSystem=strict)
excluded the helper and unit directories from ReadWritePaths, so the
unattended path could never refresh /usr/local/bin/pulse-auto-update.sh
or rewrite the units - updater fixes only reached boxes via manual
installs. Grant the sandbox write access to both directories on
purpose.
- Because the unattended path replaces the helper bash is currently
executing, stage the new helper next to its destination and swap it
in with an atomic rename only after repo configuration succeeds. A
failed download or configure now leaves the previously working helper
in place instead of rm -f'ing it out from under the enabled timer's
ExecStart.
- Delete scripts/systemd/pulse-update.{service,timer}: orphaned
reference copies that had drifted from the units install.sh actually
generates and were referenced by nothing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First candidate on the v6.2.0 line, following stable v6.1.2. Carries
the External Probes Pro feature, the three-part multi-site Proxmox
identity isolation work, cloned machine-id collapse detection,
unattended-update service recovery, the TrueNAS same-host redirect
follow, and verified telemetry outcomes. Version pins move to
6.2.0-rc.1 across the repo root, Docker bootstrap defaults, and Helm
metadata per the deployment-installability contract; stable install
pointers remain on v6.1.2 until governed promotion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The generator emitted JSON-marshaled TypeScript (quoted keys, no trailing
commas) while the committed file is prettier-formatted, so every run dirtied
frontend-modern/src/utils/selfHostedFeatureCatalog.generated.ts with a
324-line formatting-only diff. The generator now runs the frontend's own
prettier on the file after writing it, so runs at a clean checkout leave the
tree clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two bugs from #1630 that combined to take installs down silently:
1. perform_update()'s install-failed rollback branch restored the backup
but never restarted pulse.service. Since the generated
pulse-update.service gates on ExecCondition=systemctl is-active,
every later timer run was then skipped and the install stayed down
until manual intervention. Restart is now guaranteed by a
service_was_active-guarded restart in that branch plus an
ensure_service_restarted RETURN trap so no exit path can miss it
(re-fix of #1323, originally c0b3a0e66, lost in 778a2577b and only
partially restored in 672e81985).
2. install.sh aborted under errexit when writing the /bin/update helper
on a read-only filesystem - after the new binary was installed and
the service stopped, landing in bug 1's no-restart branch. The stock
pulse-update.service uses ProtectSystem=strict, so /bin and
/usr/local/bin are read-only on stock unattended updates; transient
read-only remounts hit the same path. The helper write, PATH
appends, and the /usr/local/bin/pulse symlink are now idempotent and
non-fatal with a warning (install_binary_symlink).
Contract: deployment-installability now pins fail-closed service
availability for unattended updates and non-fatal writes outside the
hardened unit's writable set, with proofs in pulse_auto_update_test.go
and root_install_sh_test.go plus shell regression coverage in
scripts/tests/test-pulse-auto-update.sh (installer-exits-nonzero path)
and scripts/tests/test-install-update-resilience.sh (read-only helper
and symlink paths, verified under set -e, root-safe via ENOTDIR).
Fixes#1630
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent installer (scripts/install.sh) had no free-space preflight, so on
RAM-rooted appliances (QNAP QTS, Unraid) it ran all the way to the download
before dying with an unhelpful ENOSPC. Lift the server installer's disk
headroom check into the agent installer: verify temp and install-dir space
(including the shared-filesystem case) before downloading and in
--preflight-only mode, with a TMPDIR hint in the failure message.
The QNAP and Unraid watchdog loops also shell-appended agent stdout to
/var/log/pulse-agent.log with no rotation, which could fill the RAM root on
its own. Pass --log-file so the agent's rotating writer engages (QNAP: data
volume state dir; Unraid: /var/log/pulse-agent with size-capped rotation),
discard the now-duplicate stdout mirror, and keep the watchdogs' own messages
in a small self-trimming log.
Document the TMPDIR override for constrained roots in docs/UNIFIED_AGENT.md.
Fixes#1617 (space half; CPU half pending reporter diagnostics)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The alerts-frontend-surface accepted proof list is asserted verbatim by
canonical_completion_guard_test.py, so registering
useAlertOverridesState.test.tsx in registry.json without updating the pinned
expectation failed the guard unit tests.
Verified by running every step of the canonical-governance workflow locally
rather than only the guard I expected to trip: status, control-plane, registry
and contract audits, the Pulse Intelligence gate schema, active-target
automated and hybrid readiness proofs, and all twelve release-control unit
test modules.
An explicit --report-ip is the user naming the primary address on a
multi-NIC host, but identityFromHost appended it after the auto-detected
interface addresses while every consumer of ResourceIdentity.IPAddresses
treats the first entry as primary, so the override never changed what
the Machines table displayed. Prepend it instead.
The install script also rejected --report-ip as an unknown argument even
though the agent supports the flag, forcing hand edits to the service
unit that a later --update run would drop. Accept the flag, render it
into the service ExecStart, persist it in connection state, and
recognise it during saved-state and arg-stream recovery so updates
preserve it.
Refs #829
Contract-Neutral: behavioral fix: user-specified report-ip leads host identity addresses and the installer passes --report-ip through; no public contract change (#829)
Keep macOS notarization mandatory for every release candidate while requiring Windows Authenticode only for stable promotion, matching the publish workflow and RC4 release packet.