Preserve the exact reviewed specialist commit and its documentation and signer controls alongside the current canonical frontier.
Change-source: pulse-maintainer
Distinguish failed verification from a tampering diagnosis, unsigned history, query errors and runtime gates. Remove live-key replacement and regeneration advice, correct default store paths, and preserve matching keys and history for isolated recovery. Bind the shipped guides to regression checks and exercise recovery distinctions with the real signer and encryption manager on synthetic temporary data.
Contract-Neutral: Audit documentation and synthetic regression controls only; no signing, storage, entitlement, API or frontend runtime behaviour changes.
Change-source: pulse-maintainer
Retain configured RAID members separately from source-native totals through collection, reports and canonical read views. Correct the healthy legacy mdadm tuple from #2369 without subtracting spares from mdstat requirements or suppressing real deficits, failures and recovery warnings. Preserve observed zero-active counts through fallback, and verify collector, wire, ingestion, health and canonical activation/resolution boundaries.
Change-source: pulse-maintainer
Adapt PR2351 to current main without losing bounded probe deadlines, retries, explicit binds or confidentiality. Inspect IPv6 wildcard socket mode so unrelated servers sharing the port cannot determine Pulse health. Apply the telemetry callback only after a persisted explicit boolean transition, including stale-disk and null-input boundaries.
Retain real HTTP/HTTPS wildcard and same-port isolation regressions, callback persistence-order tests, and the owning subsystem contracts. No dependency manifests, frontend source or telemetry payload schema change. Exact runtime proof remains dependent on the unavailable root graph; it is not represented as passed.
Change-source: pulse-maintainer
Original-source: https://github.com/rcourtman/Pulse/pull/2351
Original-commit: f50c4d6e3b
Co-authored-by: rcourtman <8825017+rcourtman@users.noreply.github.com>
Try both loopback families for IPv6 wildcard listeners without changing explicit or IPv4-only binds. Reserve each family a share of the API deadline, keep all later checks on the responding address, and retry a failed observation once with a fresh bounded budget in the telemetry background runner.
Keep genuine API, UI and asset failures visible and report only a closed timeout category, never transport details. Cover both real listener families, startup recovery, failure and privacy boundaries; synchronise the disclosure and subsystem contract.
Change-source: pulse-maintainer
Relay left public checkout on 2026-09-29: it only ever connected the
Pulse Mobile app, which retires on 31 March 2027, and existing Relay
subscribers now carry Pro entitlements. The app still sold it. Every
Community install saw a "Get alerts on your phone ... Available with
Relay and Pro plans" upgrade panel on Alerts destinations, the plan
screen offered a Relay card with "Remote web access via Relay", and the
settings section was called Remote Access although Relay never reached
the web UI.
The Alerts push panel now renders only on instances that have the relay
feature, with no upsell. The settings section, nav, header and locale
catalogs say Pulse Mobile and carry the retirement date. Community sees
only the Pro comparison card and Relay-tier licenses see none; gated
mobile features name Pro as their minimum plan. The relay feature is
labelled "Pulse Relay (Mobile Connection)" in the catalog, plan copy no
longer claims remote web access, the backend pairing diagnostics point
at Settings > Pulse Mobile, and the user docs, including PRIVACY.md,
state that Relay does not provide remote web UI access.
A provider-hosted control plane refused to start once its licence was past
expiry and the 7-day grace, because startup validated it with
ValidateLicense. So at the next restart after a 60-day evaluation ran out,
or after a paying provider's subscription lapsed, the portal went down, and
the portal's Plan tab is the only place a provider can buy or renew.
Reproduced on the walkthrough lab.
Startup now accepts an authentic, key-bound licence even when lapsed
(ValidateLicenseAllowingLapse), preferring a current licence and otherwise
the one that expired later, and logs the lapse. It unlocks nothing: client
runtimes still verify the provider licence in every lease with
ValidateLicense and drop MSP capabilities, new clients are refused with
provider_msp_license_lapsed, and the refresher still refuses a lapsed
licence. The Plan panel says the evaluation or plan ended and offers the
plans, and provider-msp status reports license_lapsed without failing so a
lapsed install can still upgrade.
The same run showed the portal printing control plane error codes such as
provider_msp_license_lapsed instead of their message. The portal now
prefers an error's message over its code, which also fixes
already_subscribed and checkout_unavailable in the Plan tab.
Metric identity lived in idx_metrics_lookup, ordered (type, id, metric,
tier, timestamp). Once a series retains history, that order gives every
series its own insertion point, so each poll's commit rewrites one index
leaf page per series through the WAL and again at checkpoint. On a
three-node homelab running v6.4.5-rc.2 that came to 330-607 KB/s of disk
writes, about 10.6 KB per ~60-byte sample and 27-50 GB a day (#1966).
idx_metrics_query_all already holds the same five columns time-major, so
one poll's samples for a resource share pages. It now carries uniqueness
and serves every read, and the metric-major lookup index is dropped. With
an hour of retained history for 480 series, 30 polls write 6,022 WAL
frames instead of 25,310 (4.2x fewer); the empty-table issue-1124 estate
drops from 36,308 frames to 20,223, and its ceiling tightens from 40,000
to 23,000.
The tier-reconciliation overlap probe now names the identity index too;
left to the planner it seeks on (tier, timestamp) alone once the
metric-major tree is gone, which TestRetainedQueryPlansUseIndexes caught.
Metric-filtered reads now scan a resource's other metrics in the window:
a single-metric 500-node query went from 45 to 59 us and the 24-hour
50-disk smart_temp batch from about 0.25 to 0.49 s. All-metric dashboard
reads are unchanged. Continuous write volume is the user-visible harm, so
the contract records that trade explicitly.
The migration reuses the crash-safe swap: databases whose identity is
already enforced by a unique index defer the rebuild to startup
maintenance and swap in one transaction, and a database with no unique
index migrates synchronously. Keeping the name idx_metrics_query_all
means a downgrade rebuilds only the lookup index and leaves no orphan.
Refs #1966
A provider could stand the MSP bundle up and evaluate it without asking
anyone, but paying still meant emailing its public key to Pulse and
copying a licence file onto the host by hand. With the licence server's
self-serve purchase path (pulse-pro license-server/provider_msp_purchase),
the control plane can do the whole thing itself.
The control plane now signs a challenge with its lease signing key to
fetch the licence its subscription entitles, validates it exactly as
startup does, requires it to bind this platform's key, and keeps it in
the data directory, because the host licence is a read-only secret
mount. Startup prefers that renewed licence while it validates, so a
paying provider keeps starting after the self-issued evaluation on the
host expires. A changed licence restarts the control plane through the
existing graceful shutdown, since the plan version is read at load by
every workspace-limit, backup and status path; client workspaces keep
running. Preflight, proof, recovery and backup accept either signed
source.
The portal gains owner and admin routes for plan state, checkout, the
Stripe billing portal and an immediate licence refresh, and the plan list
comes from the licence server so it shows only what can be bought, at
Stripe's price. Adds the msp_solo plan: 3 client workspaces, the first
paid step above the 2-workspace evaluation.
The startup-watchdog race repair only converted the capture in service_health_test.go, but startup_watchdog_test.go still polls a plain bytes.Buffer while the watchdog goroutine writes the same buffer through zerolog. go test -race therefore still reports a read/write data race in TestStartupWatchdogFiresAndNamesLastPhase and the rest-0 backend shard stays red. Capture those logs through the same mutex-guarded syncBuffer so the poll and the watchdog write are serialized.
Change-source: pulse-maintainer
TestStartupWatchdogLogsPhaseAndStack polls a bytes.Buffer while the watchdog goroutine writes the same buffer through zerolog, so 'go test -race' reports a read/write data race at service_health_test.go:134 and fails the rest-0 backend shard. Capture the log through a mutex-guarded buffer so the poll and the watchdog write are serialized.
Change-source: pulse-maintainer
pkg/server/server.go is a security-privacy runtime path, so the startup
watchdog commit needs its contract obligation and an accepted proof in the
same reviewed unit. Record the local-only diagnostic boundary in the
security-privacy contract and place the watchdog proof in the
registry-accepted pkg/server/service_health_test.go.
Change-source: pulse-maintainer
Bind the UI/API listener early so port conflicts fail fast, but srv.Serve
only starts near the end of Run. Every synchronous initialization step in
between therefore runs with a live listening socket whose accept queue fills
while nothing accepts: the reported #2129 shape (bound port, non-zero Recv-Q,
no HTTP response). A stall there is invisible because no log line is emitted
while it runs.
Track the last completed startup phase and arm an independent watchdog that
fires if serving has not begun within two minutes, logging the last phase and
every goroutine stack. This makes a recurrence diagnosable from the field log
without an operator sending SIGQUIT. The watchdog is stopped when serving
begins or Run exits and is safe to stop more than once.
Change-source: pulse-maintainer
Address the remaining verified findings from the pkg/metrics/store.go review
(#2079 finding 5 and the lower-impact nits):
- writeBatch now retries the whole transaction on a retryable lock error.
BEGIN is deferred, so contention is hit at the first INSERT or COMMIT; the
old loop retried only Begin() and skipped rows. While the one-time startup
maintenance (deferred identity-index rebuild and auto-vacuum VACUUM) is
active, an extended retry budget keeps upgrade-time history instead of
dropping it after the steady-state budget.
- migrateLegacyHostResourceType probes with an indexed EXISTS before opening a
write transaction, so an already-migrated store does not dirty the WAL on
every boot.
- QueryAll caps the per-series MetricPoint preallocation so a fine requested
step over a long range cannot reserve hundreds of thousands of slots.
Regressions: TestStoreStartupMaintenanceSignalsWriteRetryBudget,
TestStoreWriteBatchExtendsRetryDuringStartupMaintenance,
TestEstimateQueryAllBatchSeriesCapacityCapsPreallocation. The
performance-and-scalability contract records the invariants.
Change-source: pulse-maintainer
The canonical completion guard requires the metrics-store hot-path proof in
pkg/metrics/store_additional_test.go. Move the rollup-gap, upsert-spread,
busy-retry and auto-vacuum regressions there and point the
performance-and-scalability contract at that file. The runtime change and
contract section are in the preceding fix(metrics) commit.
Change-source: pulse-maintainer
Address the remaining verified findings from the pkg/metrics/store.go review
(#2079 findings 2-4 and 6):
- rollupTier no longer leaps its checkpoint over a source gap. Previously a
trailing or interspersed gap advanced the checkpoint to the rollup cutoff, so
a lower tier that backfilled the gap later had those rows stranded below the
upper checkpoint and purged by retention.
- Retry conditions now classify SQLITE_BUSY/LOCKED by driver result code. The
modernc driver reports "database is locked (5) (SQLITE_BUSY)", which the
previous exact-string comparison never matched, so a busy writer dropped its
batch without retrying.
- The metrics upsert preserves an existing rollup row's min/max spread with
COALESCE instead of overwriting it with NULL.
- migrateAutoVacuum pins the pragma and VACUUM to one connection and verifies
the persisted mode, so the one-time conversion cannot silently repeat a
full-file VACUUM on every restart.
Regressions: TestStoreRollupTierBackfillsInterspersedGap,
TestStoreWriteBatchPreservesRollupSpread, TestIsRetryableWriteError,
TestIsRetryableWriteErrorMatchesDriverBusyCode,
TestStoreAutoVacuumPersistsAcrossRestart. The performance-and-scalability
contract records the invariants.
Change-source: pulse-maintainer
The shared release preflight worker exports GITHUB_ACTIONS=true so isolated single-repository checkouts take their private-sibling skip. That flag also made the load/SLO latency guards hard-fail on a contended shared host, where the tests' own documented local-contention skip is the intended behaviour. Distinguish a real GitHub Actions run by GITHUB_RUN_ID, which only a hosted runner sets, so the worker keeps the cross-repo skip while latency overruns skip instead of failing. Hosted CI enforcement is unchanged.
Change-source: pulse-maintainer
The ingestion worker closed writeCh when it observed stopCh. A concurrent WriteWithTier that passed its stopping check, or a WriteBatchSync/WriteBatchBounded caller (which never checks stopping), could then send on the closed channel and panic the process during shutdown. Reproduced deterministically: enqueueWrite after Close panics with send on closed channel, and 8 concurrent monitoring writers racing Close panic in boundedEnqueueAndWait (store.go:1034).
Never close writeCh. Drain already-queued requests non-blockingly, then process them. Late writes land in the buffered channel and are discarded with the store rather than crashing it. Record the invariant in the performance-and-scalability contract and pin it in the accepted store proof file.
Change-source: pulse-maintainer
History requests can occupy every database connection and make cheap
statistics reads wait for seconds. Reserve one read-only connection for
current committed counts, owned and closed by the metrics store.
Keep the SQL, data model, load workload and latency budgets unchanged.
Add deterministic pool-saturation, committed-data and shutdown coverage.
First-run setup saved the configured administrator but left the router authorizer on its startup identity. This split Settings capability checks: administrator-only panels remained visible while API Access and Pulse Intelligence were denied until restart. Keep the captured authorizer aligned when setup commits the identity, and clear its prior bypass on successful development reset.
Cover the exact split with real file-backed RBAC policy and session setup, preserve outsider denial and identity replacement, and synchronise concurrent policy reads. Extension implementations still require independent compatibility review; this is not installation or release acceptance.
Change-source: pulse-maintainer
Update compress and Prometheus common independently of the grouped database proposal. Common requires x/net 0.58.0; all other selected modules and existing checksums remain unchanged. Extend real exporter wire coverage for wildcard, browser and JSON fallback requests alongside gzip, protobuf and optional zstd compatibility.
Contract-Neutral: Dependency maintenance and expanded compatibility coverage only; no Pulse API, deployment policy, toolchain or enabled exporter format changes.
Change-source: pulse-maintainer
Update the coherent six-module Prometheus dependency set without changing other selected versions. Cover Pulse metric labels, protobuf and text scrape negotiation, gzip and existing identity fallback, plus the optional zstd codec. The baseline also leaves zstd disabled, so no production registration change is needed.
Contract-Neutral: Dependency and compatibility-test maintenance only; no public API, deployment, telemetry, privacy or scrape-format behaviour changes.
Change-source: pulse-maintainer
Carry the operator include override in disk reports so server filtering does not discard selected tmpfs mounts again. Keep automatic filtering for unmarked reports and agent exclusion precedence. Cover both ingest paths, collection, forwarding and wire compatibility for #1875.
Change-source: pulse-maintainer
Hosted organisation E2E failures exposed an ambiguity between runtime identity and entitlement provisioning. Assert that Community preserves a granted multi_tenant capability without manufacturing it when absent, so fixture repairs do not wrongly require a private binary or weaken capability gating.
Change-source: pulse-maintainer
An admin-browser update read does not establish access for the monitoring token. Cover the proxied target path, configured token, denial propagation and healthy endpoint retention. Remove the unsupported Sys.Audit diagnostic hint: endpoint access is determined by Proxmox, not general inventory access.
Change-source: pulse-maintainer
Contract-Neutral: Diagnostic wording and method-comment correction plus synthetic regression coverage only; no collection logic, public API, permission policy, freshness or counter semantics change.
Return persisted planning acceptance or refusal inside the investigation turn.
Keep model judgment separate from action authority and preserve accepted action
identity across provider failures. Enforce actor/request idempotency atomically
and retain complete approval and independent verification context.
Preserve unknown disk evidence, stream whitespace and historical resolution
timestamps. Keep conversation scrolling inside its own panel. Record real-model,
disposable-lab and browser qualification with explicit population limits.
Refs #1782
Expose confined, identity-bound filesystem observations through the shared
resource pipeline so investigations can distinguish an exhausted container
mount from unrelated host capacity. Keep unavailable measurements explicit.
Isolate alert-history reads from durable writes and reuse one chronological
fold across polling. Catch up through bounded durable event IDs so simultaneous
readers do not replay every retained snapshot. Retain expired actions when
investigation outcomes move back to needs attention, and keep attached
Assistant context focused.
Record live storage diagnosis, healthy and dependency controls, approved and
rejected Docker outcomes, source-bound browser proof and exact test limits.
Missing-access continuity, VM dispatch completion and remaining Assistant
orchestration defects stay open in the redesign plan.
Live funded qualification found hidden tool results and misleading action
submission outcomes. Share the result-bearing transcript across stored chat
and product history, render the retained evidence, and distinguish captured
proposals from broker acceptance. Keep review usable while Patrol is paused.
Record Gemini route pricing and exact qualification limits. Integrate current
main and repeat browser proof for the incoming login flow. Approved/rejected
recovery remains unqualified without the development command agent.
Release qualification crashed in historical baseline SQLite ingestion during the workloads-summary seed. Preserve a metrics-only diagnostic for the same 84,000-row batch shape, with row-count and integrity checks before and after reopen, without the HTTP or reflection fixtures. This does not reproduce or fix the unexplained crash.
Validation: ten focused runs on Go 1.26.7 and one on Go 1.26.8 passed; a one-repeat race run passed. The three-repeat race run timed out at 180 seconds during its final integrity check and remains retained evidence. Omitting the final seed batch makes the row-count assertion fail.
Change-source: pulse-maintainer
Missing block I/O and container image sizes could become false evidence
for diagnosis. Preserve per-direction counter presence and measured zero
through collection, resource conversion and browser rendering. Separate
new observed history from ambiguous retained disk series without deleting
old rows or changing public metric names.
Keep partial host rates distinct and persist a newly enabled Disk I/O
column across the first preference reload.
The 500-node dashboard query triggered repeated ordinal string conversions
in the SQLite driver while matching numbered parameters. Use alphabetic
named bindings to preserve current values and shared query branches without
that allocation cost.
Cover large cached scopes with changed resource families, identities,
metric filters and windows, including IDs that resemble SQL syntax.
Refs #1928
Canonical tier reconciliation rebuilt SQL and probed absent preferred tiers
for each fallback point, regressing batch reads and allocation costs. Reuse
bounded query shapes with current bindings and snapshot-scoped absence
checks, then append consecutive points directly to their output series.
Preserve coverage and ordering semantics and verify fresh bindings after
new preferred observations arrive. Integrate current main test additions.
Execute plain retained reconciliation in one current SQLite snapshot and
reuse bounded compiled statements. Preserve per-series chronology without
a metric sort while keeping display aggregation ordering explicit.
Exact-base worker comparisons cover the prior PR benchmark failures. Full
metrics/database and focused concurrent race checks pass. Final CI and
real diagnostic outcome qualification remain open.
Keep canonical disk risk, source freshness and retained history intact when
Assistant and Patrol gather evidence. Proposal acceptance validates an action
contract and must not rewrite uncertain conclusions as established root cause.
Preserve complete subscription tool batches without exposing routing envelopes
as answers. Keep wide answer tables readable and keyboard-scrollable on mobile.
Optimize retained tier reconciliation without discarding gaps or newer samples.
Record failed real-model diagnoses and outstanding autonomous qualification
separately from passing data-path and interface checks.
Non-empty aggregate tiers hid recent raw samples and missing metric series.
Unify single and batch reads with indexed overlap resolution before
downsampling, and preserve recorded extrema through subsequent rollups.
Use canonical Proxmox storage coordinates in summaries and node history.
Discovery routing does not establish an installed Agent or agent history.
Record live evidence freshness and the remaining diagnosis qualification gaps.
A provider policy refusal was classified as a connection failure, while
summary tools mixed high-utilisation heuristics with unchecked health claims.
Preserve explicit refusals before tool recovery and provide retained metrics
with source scope, observation timestamps and bucket extrema for diagnosis.
Record the live qualification limits and the shared temporal tier-query gap.
New-finding telemetry lost older activity once the run history reached its
100-entry cap. Persist a bounded daily finding tally with a separate
upgrade cursor, preserving run counts while backfilling retained findings.
Cover restart, repeated saves, upgrade, read failure and UTC-day retention.
Record the measurement boundary and retire the resolved coverage gap.
Reproduced a false resolved webhook when HTTP-success status contained memory total but omitted used. Require present CPU and memory measurements before marking node metrics available, retaining genuine zero and unrelated-field compatibility. Extend client and real-poller webhook regression coverage and monitoring contract.
Change-source: pulse-maintainer
A successful response containing null or omitted data decoded into zero-valued node metrics. Repeated polls could therefore clear an active memory incident without any usable recovery evidence. Decode the status through a pointer and reject absent data so existing unavailable-metric handling preserves the incident.
Reproduced the failure through the real poller and notification queue with a local webhook. Added absent-envelope client cases and extended lifecycle coverage to assert unavailable metrics, stable incident identity, and genuine recovery. Focused client and monitoring tests pass three repetitions under the race detector; this is not installed or off-host qualification.
Change-source: pulse-maintainer
A gateway body quoting API error 403 must not discard cached backups. Use the client's typed response status before legacy text fallback; cover gateway failures and genuine terminal responses.
Change-source: pulse-maintainer
Limited tokens intentionally omit node metrics, but later gateway failures must remain visible rather than being classified from permission text. Exercise restriction, outage and recovery on the same client to protect this monitoring boundary.
Change-source: pulse-maintainer
The persistence profile labelled /proc/self/io write_bytes as physical writes, which could mislead release qualification into attributing process accounting to device wear. Report process writes and cancellation separately, preserving unavailable counters as unknown rather than implying zero.
Add focused accounting coverage for malformed, missing, overflowing and decreasing counters. This is diagnostic-only: it neither changes runtime persistence nor establishes installed write cost or release readiness.
Validation: focused TestIssue1124ProcessIOAccounting and serialised persistence profile passed before this message-only repair; git diff --check passed. The tested tree is unchanged.
Change-source: pulse-maintainer
Issue #1895 reports parity alerts when mdNumDisks=0 on a pool-only Unraid system. Array service state alone does not establish that a parity array exists.
Preserve the optional disk count from collection through canonical runtime conversion and suppress only the no-parity warning for an explicit zero. Missing or malformed counts retain legacy behaviour, and disk failure reasons remain active.
Validated focused Unraid tests in hostagent, storagehealth, monitoring, unifiedresources and alerts, including JSON zero preservation and canonical round trip. The new pool-only regression fails against the previous warning condition. Both agent and server need this change; no release or reporter retest is claimed.
Change-source: pulse-maintainer
A three-week decline in investigations/new_findings (7.2% to 5.0%) looked
like a Patrol regression. It is not one. The ratio is not a rate at all:
the two counters come from different stores, cover different spans, and
are drawn from populations that barely overlap.
new_findings_30d sums run.NewFindings over history.Runs, which
SavePatrolRunHistory caps at MaxPatrolRunHistory, so it covers at most the
last hundred runs rather than thirty days. investigations_30d instead
scans the current findings store and counts surviving finding records
investigated in-window, including findings created before it, which is how
the paid cohort read 128.57% in the week to 2026-08-25. Finding
ShouldInvestigate returns false at monitor autonomy and effective autonomy
is licence-gated, so free installs produced 4384 findings and 1
investigation while 67 paid installs produced 242.
The decline was composition: flat in version-stable installs, and fleet
investigations rose once the single install that swung the total by 38 was
excluded. Finding-detection code is identical between v6.3.2 and v6.4.1.
The new test pins the asymmetry behind the bad denominator. Its twin
already asserts that runs_30d ignores the history cap after 63c40ebe5e;
nothing asserted that the findings loop immediately below it does not, so
the truncation could regress or be mistaken for a thirty-day total
unnoticed. Fixing it needs a per-day findings tally alongside DailyRuns,
which the coverage gap tracks as its own slice.
pkg/server tests boot the real server through Run() with the version
literal "test-version", which internal/updates normalizes to
0.0.0-test-version. Each test runs against its own t.TempDir(), so every
run minted a fresh install ID. The startup ping waits two minutes and so
never fired inside a short test, but the service-health failure reporter
added on 2026-08-29 sends synchronously from a deferred handler as soon
as Run() returns an error, so every CI shard containing pkg/server posted
one ping.
The licence server recorded 317 single-ping installs between 2026-08-29
and 2026-09-03 - 311 from linux/amd64 CI runners, 3 from a maintainer
workstation - still arriving at roughly 60 a day. The canonical clean
denominator excludes single-ping installs and was unaffected, but raw
install counts and the operator-evidence blocked-cause read counted them
as real installations.
A test binary is not an installation, which is the same reason mock mode
already suppresses pings, so the guard belongs beside it in the telemetry
package rather than at the four call sites: send() now refuses the
production endpoint whenever testing.Testing() reports true. The check
compares against productionPingEndpoint, so telemetry's own tests keep
asserting on real ping content through a redirected endpoint, and the
server tests additionally opt out at the config layer to say so locally.
Incorporate the metrics startup hook capture from PR #1868 while retaining
the reviewed causal cleanup and barrier-ordering coverage. This advances the
open publication proposal without rewriting any accepted commit.
Change-source: pulse-maintainer
Contract-Neutral: Integration reconciliation only; no additional public contract delta.
TestNewStoreDefersStartupMaintenance bounded NewStore at 200ms. On a
slow CI disk that bound tripped, the test failed, and its NewStore
goroutine kept running: the maintenance worker it spawned read the
package-level startupMaintenanceHook after the next test had installed
its own closure, closed that test's started channel a second time, and
panicked the whole rest-1 shard (run 33630289317, attempt 1).
Capture the hook once in NewStore so a store can only ever call the
hook that was installed when it was built. Prove deferral by ordering
instead of wall clock: the hook parks the worker, and NewStore must
return while it is parked. A regression to inline maintenance now
blocks that receive until the package timeout instead of flaking.
Cleanup releases the worker and waits for NewStore before restoring
the hook, so a failed run cannot leak a parked store either.
Contract-Neutral: behavioral test-flake fix with no public contract delta
Retain the exact core-runtime candidate commit and integrate its deterministic maintenance-worker ordering and cleanup coverage.
Change-source: pulse-maintainer
The race-enabled suite can take longer than the tests' fixed sleeps while opening SQLite stores. Assert worker ordering through channels instead, and always release and join blocked maintenance workers so a failed assertion cannot contaminate the following test.
Change-source: pulse-maintainer