mirror of
https://github.com/rcourtman/Pulse.git
synced 2026-10-03 12:47:49 +00:00
Add release service health telemetry
This commit is contained in:
parent
db7e26deac
commit
b75a5aeec2
17 changed files with 1064 additions and 59 deletions
|
|
@ -16,7 +16,7 @@ third-party analytics, support diagnostics, or ordinary Settings surfaces.
|
|||
|
||||
Pulse includes outbound usage telemetry that is **enabled by default**. It sends a lightweight ping on startup and once every 24 hours with a rotating pseudonymous install ID to help me understand how many active installations exist, which releases are actually deployed, which features are in use, and whether Patrol control and governed Pulse Intelligence operations are being adopted.
|
||||
|
||||
The telemetry payload does not include hostnames, credentials, infrastructure identifiers, IP addresses, URLs, paths, locale, prompts, chat messages, command text, action output, token values, names, email addresses, or account identifiers. Lifecycle and outcome signals are deliberately limited to closed buckets, booleans, and aggregate counts. Pulse does not send browser events or an event-level clickstream. See the full field list below.
|
||||
The telemetry payload does not include hostnames, credentials, infrastructure identifiers, IP addresses, URLs, paths, locale, prompts, chat messages, command text, action output, token values, names, email addresses, or account identifiers. Lifecycle and outcome signals are deliberately limited to closed buckets, booleans, and aggregate counts. Local service-health signals are likewise limited to closed buckets, booleans, and normalized release versions. Pulse does not send browser events or an event-level clickstream. The local service-health probe requests only Pulse's own bound listener and never transmits the listener address, request URL, response content, or raw failure text. See the full field list below.
|
||||
|
||||
While mock/demo fixture mode is enabled, Pulse suppresses outbound telemetry entirely: a mock-mode instance reports a synthetic fixture fleet rather than a real installation, so it never pings.
|
||||
|
||||
|
|
@ -37,7 +37,7 @@ Every field is listed below with the reason it exists. Nothing else is included
|
|||
|
||||
| Field | Example | Purpose |
|
||||
|-------|---------|---------|
|
||||
| Schema version | `3` | Identify the exact payload contract so old and new signals are not mixed silently |
|
||||
| Schema version | `13` | Identify the exact payload contract so old and new signals are not mixed silently |
|
||||
| Sent at | `2026-07-23T08:30:00Z` | Date the individual heartbeat without sending a history of client activity |
|
||||
| Install ID | `a1b2c3d4-...` | Distinguish active installations within one rotation window without tying telemetry to an account or person |
|
||||
| Version | `6.0.0-rc.1` | Track the canonical release identity currently deployed |
|
||||
|
|
@ -120,6 +120,13 @@ Every field is listed below with the reason it exists. Nothing else is included
|
|||
| Update successes 30d | `1` | Count successful update attempts in the current 30-day telemetry window |
|
||||
| Update failures 30d | `1` | Count failed or rolled-back update attempts in the current 30-day telemetry window without sending raw errors, logs, URLs, or command output |
|
||||
| Update last failure category | `download` | Send only a coarse category for the latest update failure, such as `download`, `signature`, `checksum`, `disk_space`, `extract`, `backup`, `apply`, `restart`, `rolled_back`, or `unknown` |
|
||||
| Service health observed | `true`/`false` | Distinguish a release that performed the bounded local UI/API self-check from an older release with no signal |
|
||||
| Service health healthy | `true`/`false` | Report whether Pulse's locally bound listener served a healthy API response, UI document, and every referenced local frontend asset without sending an address, URL, response, or error text |
|
||||
| Service health failure category | `listener`, `startup`, `runtime`, `api_connectivity`, `api_status`, `ui_status`, `frontend_assets`, or `unknown` | Classify a failed local self-check or startup path into one fixed category without sending the listener address, request URL, HTTP body, asset name, IP address, or raw error |
|
||||
| Service health cohort | `first_observation`, `same_version`, or `version_change` | Mark whether this direct observation is the first schema-v13 observation, another observation of the same release, or the first observed state after a release-version change |
|
||||
| Service health previous version | `6.4.0` | Retain only the normalized immediately previous Pulse release identity so aggregate reporting can compare a post-upgrade before/after cohort without treating rolling counters as new-release activity |
|
||||
| Service health previous observed | `true`/`false` | State whether the immediately previous release recorded the bounded local service-health signal |
|
||||
| Service health previous healthy | `true`/`false` | State whether that immediately previous release's last bounded local observation was healthy, without retaining its address, URL, response, asset names, or failure text |
|
||||
| Node test attempts 30d | `3` | Count node connection tests that reached the connection stage in the current 30-day telemetry window, without sending hostnames, addresses, credentials, or error text |
|
||||
| Node test failures 30d | `2` | Count those connection tests that could not reach or authenticate against the target, so an install that tried to add a node and failed is distinguishable from one that never attempted it, without sending hostnames, addresses, credentials, or error text |
|
||||
| Pulse Intelligence loop configured | `true`/`false` | See whether Assistant, Patrol, governed actions, or external-agent access is configured so adoption can be measured without sending configuration details |
|
||||
|
|
@ -219,6 +226,15 @@ configuration, destination rejection, and unknown failures. Classification is
|
|||
performed locally; raw error text and destination/provider identity are never
|
||||
included in the payload.
|
||||
|
||||
Telemetry schema v13 adds a direct local UI/API service observation. Pulse
|
||||
checks its own bound listener, `/api/health`, the UI document, and the local
|
||||
frontend assets referenced by that document. It retains only the current and
|
||||
immediately previous normalized release observation, so a version-change
|
||||
cohort can compare before and after health without attributing rolling update
|
||||
or usage counters to the new release. Listener, startup, runtime, API, UI, and
|
||||
asset failures collapse into fixed categories. Addresses, URLs, IP addresses,
|
||||
asset names, response bodies, and raw errors are neither sent nor stored.
|
||||
|
||||
#### Server-side handling and retention
|
||||
|
||||
- Telemetry pings are stored on the Pulse license server only for aggregate install/use analysis.
|
||||
|
|
@ -226,6 +242,7 @@ included in the payload.
|
|||
- Pulse may derive aggregate Pulse Intelligence adoption reports from those same rows, including whether an install reached Patrol issue activity, Patrol resolution, Assistant, direct external-agent, or MCP collaboration, Patrol mode starter use, paid Patrol mode cohorts, governed-action activity, approved or rejected action decisions, approved action success, completed Patrol control work, recent retention, and observed free-to-paid movement within the source window. Those reports do not add prompts, findings, resource identifiers, tool names, tool inputs, tool outputs, command payloads, action outputs, account links, or exact commercial tiers.
|
||||
- Aggregate reports preserve the closed deployment-method buckets, but treat them as best-effort current-runtime evidence. In particular, `container_other` and `binary_other` are unknown fallbacks for many upgraded installs, not precise original installation provenance.
|
||||
- Aggregate reports describe the known-age bucket as time since Pulse first created the schema-v2 lifecycle record. It is only a lower bound for an upgraded install and must not be presented as original installation age.
|
||||
- Aggregate release-health reports use the direct schema-v13 current and immediately previous observations. They do not infer new-release health from rolling update, feature, alert, or action counters.
|
||||
- External-agent/MCP activity is stored only as a coarse adapter-origin flag plus capability-class counters: context, event stream, provisioning, operator state, findings, and action requests.
|
||||
- The receiver stores only fields in its versioned telemetry allowlist. A cross-repository parity check prevents client fields from being silently dropped and prevents the storage contract from growing beyond the disclosed payload.
|
||||
- Telemetry rows older than **90 days** are purged automatically.
|
||||
|
|
|
|||
|
|
@ -9685,6 +9685,21 @@ the operator preview cannot silently omit fields that the sender will export.
|
|||
The frontend keeps those additions optional only where the Go wire field is
|
||||
optional; numeric counters remain required and default to zero in fixtures.
|
||||
|
||||
### Telemetry payload parity spans three surfaces at schema v13
|
||||
|
||||
Schema v13 preserves the same sender, receiver, and Settings-preview parity
|
||||
rule while adding the direct local release service-health observation. The
|
||||
preview must expose the observed and healthy booleans, fixed failure category,
|
||||
fixed first/same/version-change cohort, normalized immediately previous
|
||||
release, and its observed and healthy booleans exactly as the Go sender would
|
||||
transmit them. Optional failure, cohort, and previous-version strings follow
|
||||
the Go JSON omission contract; observation booleans remain required. Neither
|
||||
the API response nor the TypeScript interface may add a listener address, URL,
|
||||
IP address, asset path, response content, raw error, or account/customer
|
||||
identity. The schema parity checker remains the executable cross-repository
|
||||
proof that the browser preview and private allowlisted receiver cannot drift
|
||||
from the public runtime payload.
|
||||
|
||||
### Per-tenant resource stores are released on offboarding and shutdown
|
||||
|
||||
`ResourceHandlers.getStore` opens a SQLite handle per org and caches it for the
|
||||
|
|
|
|||
|
|
@ -7129,6 +7129,7 @@
|
|||
"internal/crypto/crypto.go",
|
||||
"internal/logging/logging.go",
|
||||
"internal/securityutil/secure_storage_dir.go",
|
||||
"internal/telemetry/service_health.go",
|
||||
"internal/telemetry/telemetry.go",
|
||||
"pkg/audit/async_logger.go",
|
||||
"pkg/audit/audit.go",
|
||||
|
|
@ -7138,6 +7139,7 @@
|
|||
"pkg/auth/sqlite_manager.go",
|
||||
"pkg/extensions/audit_admin.go",
|
||||
"pkg/server/server.go",
|
||||
"pkg/server/service_health.go",
|
||||
"pkg/server/telemetry_pulse_intelligence.go",
|
||||
"pkg/tlsutil/fingerprint.go",
|
||||
"scripts/telemetry_adoption_report.py",
|
||||
|
|
@ -7161,7 +7163,9 @@
|
|||
"internal/cloudcp/auth/magiclink_test.go",
|
||||
"internal/config/config_load_test.go",
|
||||
"internal/config/watcher_test.go",
|
||||
"internal/telemetry/service_health_test.go",
|
||||
"internal/telemetry/telemetry_test.go",
|
||||
"pkg/server/service_health_test.go",
|
||||
"pkg/server/telemetry_pulse_intelligence_test.go",
|
||||
"pkg/tlsutil/tlsutil_test.go",
|
||||
"scripts/tests/test_telemetry_adoption_report.py"
|
||||
|
|
@ -7176,7 +7180,10 @@
|
|||
"docs/PRIVACY.md",
|
||||
"frontend-modern/public/docs/PRIVACY.md",
|
||||
"frontend-modern/src/components/Settings/useSystemSettingsState.ts",
|
||||
"internal/telemetry/service_health.go",
|
||||
"internal/telemetry/telemetry.go",
|
||||
"pkg/server/server.go",
|
||||
"pkg/server/service_health.go",
|
||||
"pkg/server/telemetry_pulse_intelligence.go",
|
||||
"scripts/telemetry_adoption_report.py",
|
||||
"SECURITY.md"
|
||||
|
|
@ -7186,7 +7193,9 @@
|
|||
"exact_files": [
|
||||
"frontend-modern/src/stores/__tests__/systemSettings.test.ts",
|
||||
"internal/api/system_settings_telemetry_test.go",
|
||||
"internal/telemetry/service_health_test.go",
|
||||
"internal/telemetry/telemetry_test.go",
|
||||
"pkg/server/service_health_test.go",
|
||||
"pkg/server/telemetry_pulse_intelligence_test.go",
|
||||
"scripts/tests/test_telemetry_adoption_report.py"
|
||||
]
|
||||
|
|
|
|||
|
|
@ -106,28 +106,30 @@ promise visible inside the product cannot drift from the repository policy.
|
|||
26. `internal/config/config.go`
|
||||
27. `internal/config/watcher.go`
|
||||
28. `internal/telemetry/telemetry.go`
|
||||
29. `pkg/server/telemetry_pulse_intelligence.go`
|
||||
30. `internal/api/router_routes_auth_security.go`
|
||||
31. `internal/crypto/crypto.go`
|
||||
32. `internal/securityutil/secure_storage_dir.go`
|
||||
33. `internal/cloudcp/auth/magiclink.go`
|
||||
34. `internal/cloudcp/auth/magiclink_store.go`
|
||||
35. `pkg/tlsutil/fingerprint.go`
|
||||
36. `pkg/audit/audit.go`
|
||||
37. `pkg/audit/async_logger.go`
|
||||
38. `pkg/audit/sqlite_logger.go`
|
||||
39. `pkg/audit/signer.go`
|
||||
40. `pkg/audit/sqlite_factory.go`
|
||||
41. `pkg/extensions/audit_admin.go`
|
||||
42. `scripts/telemetry_adoption_report.py`
|
||||
43. `frontend-modern/src/components/Settings/DataHandlingPanel.tsx`
|
||||
44. `frontend-modern/src/components/Settings/dataHandlingPanelModel.ts`
|
||||
45. `internal/api/agent_exec_token_binding.go`
|
||||
46. `internal/logging/logging.go`
|
||||
47. `pkg/auth/rbac.go`
|
||||
48. `pkg/auth/rbac_manager.go`
|
||||
49. `pkg/auth/sqlite_manager.go`
|
||||
50. `pkg/server/server.go`
|
||||
29. `internal/telemetry/service_health.go`
|
||||
30. `pkg/server/service_health.go`
|
||||
31. `pkg/server/telemetry_pulse_intelligence.go`
|
||||
32. `internal/api/router_routes_auth_security.go`
|
||||
33. `internal/crypto/crypto.go`
|
||||
34. `internal/securityutil/secure_storage_dir.go`
|
||||
35. `internal/cloudcp/auth/magiclink.go`
|
||||
36. `internal/cloudcp/auth/magiclink_store.go`
|
||||
37. `pkg/tlsutil/fingerprint.go`
|
||||
38. `pkg/audit/audit.go`
|
||||
39. `pkg/audit/async_logger.go`
|
||||
40. `pkg/audit/sqlite_logger.go`
|
||||
41. `pkg/audit/signer.go`
|
||||
42. `pkg/audit/sqlite_factory.go`
|
||||
43. `pkg/extensions/audit_admin.go`
|
||||
44. `scripts/telemetry_adoption_report.py`
|
||||
45. `frontend-modern/src/components/Settings/DataHandlingPanel.tsx`
|
||||
46. `frontend-modern/src/components/Settings/dataHandlingPanelModel.ts`
|
||||
47. `internal/api/agent_exec_token_binding.go`
|
||||
48. `internal/logging/logging.go`
|
||||
49. `pkg/auth/rbac.go`
|
||||
50. `pkg/auth/rbac_manager.go`
|
||||
51. `pkg/auth/sqlite_manager.go`
|
||||
52. `pkg/server/server.go`
|
||||
|
||||
## Shared Boundaries
|
||||
|
||||
|
|
@ -436,6 +438,14 @@ the `white_label` branding entitlement.
|
|||
credentials, recipient details, alert evidence, or other tenant-private
|
||||
notification configuration into the resource or incident timeline.
|
||||
6. Change operator-facing telemetry/adoption reporting through `scripts/telemetry_adoption_report.py` together with the privacy disclosure whenever release-identity interpretation changes.
|
||||
Release service-health telemetry must come from a bounded loopback probe of
|
||||
the listener Pulse actually bound. It may report only whether the API, UI,
|
||||
and referenced frontend assets were served, a fixed failure category, and
|
||||
the immediately previous normalized release observation. It must not report
|
||||
listener addresses, URLs, asset names, response bodies, errors, hostnames,
|
||||
customer identity, or account identity. Adoption reporting must interpret
|
||||
those direct current/previous observations as release cohorts rather than
|
||||
attributing rolling historical counters to the current release.
|
||||
The adoption report excludes mock-fixture-fleet-signature rows (120×N
|
||||
Kubernetes pods with 7×N VMware hosts, the `internal/mock` template) from
|
||||
adoption reads by default and must disclose the excluded row/install
|
||||
|
|
@ -1205,6 +1215,18 @@ identify only the governed class (`download`, `signature`, `checksum`,
|
|||
`cancelled`, or `unknown`). It must not export raw updater error text,
|
||||
download URLs, command output, log lines, paths, hostnames, release asset URLs,
|
||||
checksums, signatures, or operator-entered values.
|
||||
Schema v13 adds a direct local service-health observation so a process-level
|
||||
telemetry heartbeat is not mistaken for proof that the installed UI and API
|
||||
are being served. The runtime probes its bound listener through loopback and
|
||||
reports only observed/healthy booleans, one fixed failure class (`listener`,
|
||||
`startup`, `runtime`, `api_connectivity`, `api_status`, `ui_status`,
|
||||
`frontend_assets`, or `unknown`), a fixed observation cohort, and the
|
||||
immediately previous normalized release's observed/healthy booleans. No probe
|
||||
target, listener address, URL, IP address, asset path, response content, raw
|
||||
error, account, customer, or infrastructure identity may enter the payload or
|
||||
persisted receiver row. The previous-release fields are direct adjacent-release
|
||||
observations, not 30-day update counters, and are the only valid basis for a
|
||||
before/after release-health cohort in the adoption report.
|
||||
That same outbound usage telemetry floor now also permits only content-free Pulse
|
||||
Patrol control and governed Pulse Intelligence operations adoption flags and
|
||||
counters inside the same rotating 30-day telemetry window:
|
||||
|
|
|
|||
|
|
@ -1,23 +1,19 @@
|
|||
{
|
||||
"version": 1,
|
||||
"base_sha": "ea3f6388f3eb00381b02f33a78bca403dafa6dd7",
|
||||
"verified_at": "2026-08-29T08:24:25Z",
|
||||
"base_sha": "db7e26deac2a77dd5eff1dfb5bf2f1546683d5d6",
|
||||
"verified_at": "2026-08-29T10:03:51Z",
|
||||
"result": "passed",
|
||||
"changed_paths": [
|
||||
"frontend-modern/src/hooks/useUnifiedResources.ts",
|
||||
"frontend-modern/src/hooks/useWorkloads.ts",
|
||||
"frontend-modern/src/types/resource.ts"
|
||||
"frontend-modern/src/api/settings.ts"
|
||||
],
|
||||
"content_sha256": {
|
||||
"frontend-modern/src/hooks/useUnifiedResources.ts": "9a30ba0035976f356abad5b9803bbddfdbafe28a77ef9d461c7ed3f8e70c7595",
|
||||
"frontend-modern/src/hooks/useWorkloads.ts": "405e74ddd3cb490b54661fe6cd2d603adf47a56b1737f99c384f6a27cce2446c",
|
||||
"frontend-modern/src/types/resource.ts": "0effa6ccaa032a106e9194d3f2c4665061aeb72abe31b062c0c8a5f002a7f5c1"
|
||||
"frontend-modern/src/api/settings.ts": "e54e7adfb7c66c8626e2cb0abcf20c1c7903b3d26e2f0050750a97a89811eab2"
|
||||
},
|
||||
"routes": ["/proxmox/overview", "/proxmox/backups/date"],
|
||||
"routes": ["/settings/general"],
|
||||
"viewports": [
|
||||
{
|
||||
"width": 1600,
|
||||
"height": 1000
|
||||
"width": 1280,
|
||||
"height": 800
|
||||
},
|
||||
{
|
||||
"width": 390,
|
||||
|
|
@ -25,15 +21,14 @@
|
|||
}
|
||||
],
|
||||
"states": [
|
||||
"mock-backed Proxmox Overview first paint with completed, running, and genuinely absent backup states plus populated guest uptime",
|
||||
"full-page Overview refresh retaining backup and uptime evidence without an interim unknown state",
|
||||
"Overview revisited after navigating through the Proxmox backup workflow",
|
||||
"desktop and narrow layouts contained within the viewport without horizontal document overflow"
|
||||
"authenticated General settings with outbound telemetry locked off by PULSE_TELEMETRY",
|
||||
"current heartbeat preview rendered from the schema-v13 backend at desktop width",
|
||||
"current heartbeat preview rendered from the schema-v13 backend at narrow width"
|
||||
],
|
||||
"interactions": [
|
||||
"opened Proxmox Overview at 1600x1000 and confirmed 140 rendered backup indicators included completed, running, and no-backup states with no unknown uptime",
|
||||
"refreshed Overview at 1600x1000 and confirmed the same 140 backup indicators and populated uptime remained on first paint",
|
||||
"navigated to Proxmox Backups By date and back to Overview at 1600x1000, then confirmed backup and uptime evidence remained complete",
|
||||
"repeated direct load, full-page refresh, and Backups-to-Overview route switching at 390x844; each state rendered 36 backup indicators, populated uptime, and no horizontal overflow"
|
||||
"opened General settings and selected Preview payload",
|
||||
"confirmed the preview explains that telemetry is disabled while still showing the exact payload Pulse would send",
|
||||
"confirmed schema_version 13 plus service_health_observed, service_health_healthy, and service_health_cohort in the rendered payload at 1280x800",
|
||||
"confirmed the same service-health payload fields remained visible and readable at 390x844"
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -16,7 +16,7 @@ third-party analytics, support diagnostics, or ordinary Settings surfaces.
|
|||
|
||||
Pulse includes outbound usage telemetry that is **enabled by default**. It sends a lightweight ping on startup and once every 24 hours with a rotating pseudonymous install ID to help me understand how many active installations exist, which releases are actually deployed, which features are in use, and whether Patrol control and governed Pulse Intelligence operations are being adopted.
|
||||
|
||||
The telemetry payload does not include hostnames, credentials, infrastructure identifiers, IP addresses, URLs, paths, locale, prompts, chat messages, command text, action output, token values, names, email addresses, or account identifiers. Lifecycle and outcome signals are deliberately limited to closed buckets, booleans, and aggregate counts. Pulse does not send browser events or an event-level clickstream. See the full field list below.
|
||||
The telemetry payload does not include hostnames, credentials, infrastructure identifiers, IP addresses, URLs, paths, locale, prompts, chat messages, command text, action output, token values, names, email addresses, or account identifiers. Lifecycle and outcome signals are deliberately limited to closed buckets, booleans, and aggregate counts. Local service-health signals are likewise limited to closed buckets, booleans, and normalized release versions. Pulse does not send browser events or an event-level clickstream. The local service-health probe requests only Pulse's own bound listener and never transmits the listener address, request URL, response content, or raw failure text. See the full field list below.
|
||||
|
||||
While mock/demo fixture mode is enabled, Pulse suppresses outbound telemetry entirely: a mock-mode instance reports a synthetic fixture fleet rather than a real installation, so it never pings.
|
||||
|
||||
|
|
@ -37,7 +37,7 @@ Every field is listed below with the reason it exists. Nothing else is included
|
|||
|
||||
| Field | Example | Purpose |
|
||||
|-------|---------|---------|
|
||||
| Schema version | `3` | Identify the exact payload contract so old and new signals are not mixed silently |
|
||||
| Schema version | `13` | Identify the exact payload contract so old and new signals are not mixed silently |
|
||||
| Sent at | `2026-07-23T08:30:00Z` | Date the individual heartbeat without sending a history of client activity |
|
||||
| Install ID | `a1b2c3d4-...` | Distinguish active installations within one rotation window without tying telemetry to an account or person |
|
||||
| Version | `6.0.0-rc.1` | Track the canonical release identity currently deployed |
|
||||
|
|
@ -120,6 +120,13 @@ Every field is listed below with the reason it exists. Nothing else is included
|
|||
| Update successes 30d | `1` | Count successful update attempts in the current 30-day telemetry window |
|
||||
| Update failures 30d | `1` | Count failed or rolled-back update attempts in the current 30-day telemetry window without sending raw errors, logs, URLs, or command output |
|
||||
| Update last failure category | `download` | Send only a coarse category for the latest update failure, such as `download`, `signature`, `checksum`, `disk_space`, `extract`, `backup`, `apply`, `restart`, `rolled_back`, or `unknown` |
|
||||
| Service health observed | `true`/`false` | Distinguish a release that performed the bounded local UI/API self-check from an older release with no signal |
|
||||
| Service health healthy | `true`/`false` | Report whether Pulse's locally bound listener served a healthy API response, UI document, and every referenced local frontend asset without sending an address, URL, response, or error text |
|
||||
| Service health failure category | `listener`, `startup`, `runtime`, `api_connectivity`, `api_status`, `ui_status`, `frontend_assets`, or `unknown` | Classify a failed local self-check or startup path into one fixed category without sending the listener address, request URL, HTTP body, asset name, IP address, or raw error |
|
||||
| Service health cohort | `first_observation`, `same_version`, or `version_change` | Mark whether this direct observation is the first schema-v13 observation, another observation of the same release, or the first observed state after a release-version change |
|
||||
| Service health previous version | `6.4.0` | Retain only the normalized immediately previous Pulse release identity so aggregate reporting can compare a post-upgrade before/after cohort without treating rolling counters as new-release activity |
|
||||
| Service health previous observed | `true`/`false` | State whether the immediately previous release recorded the bounded local service-health signal |
|
||||
| Service health previous healthy | `true`/`false` | State whether that immediately previous release's last bounded local observation was healthy, without retaining its address, URL, response, asset names, or failure text |
|
||||
| Node test attempts 30d | `3` | Count node connection tests that reached the connection stage in the current 30-day telemetry window, without sending hostnames, addresses, credentials, or error text |
|
||||
| Node test failures 30d | `2` | Count those connection tests that could not reach or authenticate against the target, so an install that tried to add a node and failed is distinguishable from one that never attempted it, without sending hostnames, addresses, credentials, or error text |
|
||||
| Pulse Intelligence loop configured | `true`/`false` | See whether Assistant, Patrol, governed actions, or external-agent access is configured so adoption can be measured without sending configuration details |
|
||||
|
|
@ -219,6 +226,15 @@ configuration, destination rejection, and unknown failures. Classification is
|
|||
performed locally; raw error text and destination/provider identity are never
|
||||
included in the payload.
|
||||
|
||||
Telemetry schema v13 adds a direct local UI/API service observation. Pulse
|
||||
checks its own bound listener, `/api/health`, the UI document, and the local
|
||||
frontend assets referenced by that document. It retains only the current and
|
||||
immediately previous normalized release observation, so a version-change
|
||||
cohort can compare before and after health without attributing rolling update
|
||||
or usage counters to the new release. Listener, startup, runtime, API, UI, and
|
||||
asset failures collapse into fixed categories. Addresses, URLs, IP addresses,
|
||||
asset names, response bodies, and raw errors are neither sent nor stored.
|
||||
|
||||
#### Server-side handling and retention
|
||||
|
||||
- Telemetry pings are stored on the Pulse license server only for aggregate install/use analysis.
|
||||
|
|
@ -226,6 +242,7 @@ included in the payload.
|
|||
- Pulse may derive aggregate Pulse Intelligence adoption reports from those same rows, including whether an install reached Patrol issue activity, Patrol resolution, Assistant, direct external-agent, or MCP collaboration, Patrol mode starter use, paid Patrol mode cohorts, governed-action activity, approved or rejected action decisions, approved action success, completed Patrol control work, recent retention, and observed free-to-paid movement within the source window. Those reports do not add prompts, findings, resource identifiers, tool names, tool inputs, tool outputs, command payloads, action outputs, account links, or exact commercial tiers.
|
||||
- Aggregate reports preserve the closed deployment-method buckets, but treat them as best-effort current-runtime evidence. In particular, `container_other` and `binary_other` are unknown fallbacks for many upgraded installs, not precise original installation provenance.
|
||||
- Aggregate reports describe the known-age bucket as time since Pulse first created the schema-v2 lifecycle record. It is only a lower bound for an upgraded install and must not be presented as original installation age.
|
||||
- Aggregate release-health reports use the direct schema-v13 current and immediately previous observations. They do not infer new-release health from rolling update, feature, alert, or action counters.
|
||||
- External-agent/MCP activity is stored only as a coarse adapter-origin flag plus capability-class counters: context, event stream, provisioning, operator state, findings, and action requests.
|
||||
- The receiver stores only fields in its versioned telemetry allowlist. A cross-repository parity check prevents client fields from being silently dropped and prevents the storage contract from growing beyond the disclosed payload.
|
||||
- Telemetry rows older than **90 days** are purged automatically.
|
||||
|
|
|
|||
|
|
@ -75,6 +75,13 @@ const mockTelemetryPreviewPayload = {
|
|||
update_successes_30d: 0,
|
||||
update_failures_30d: 0,
|
||||
update_last_failure_category: undefined,
|
||||
service_health_observed: true,
|
||||
service_health_healthy: true,
|
||||
service_health_failure_category: undefined,
|
||||
service_health_cohort: 'same_version',
|
||||
service_health_previous_version: '6.3.0',
|
||||
service_health_previous_observed: true,
|
||||
service_health_previous_healthy: true,
|
||||
node_test_attempts_30d: 0,
|
||||
node_test_failures_30d: 0,
|
||||
alerts_fired_30d: 0,
|
||||
|
|
|
|||
|
|
@ -76,6 +76,13 @@ export interface TelemetryPingPreview {
|
|||
update_successes_30d: number;
|
||||
update_failures_30d: number;
|
||||
update_last_failure_category?: string;
|
||||
service_health_observed: boolean;
|
||||
service_health_healthy: boolean;
|
||||
service_health_failure_category?: string;
|
||||
service_health_cohort?: string;
|
||||
service_health_previous_version?: string;
|
||||
service_health_previous_observed: boolean;
|
||||
service_health_previous_healthy: boolean;
|
||||
node_test_attempts_30d: number;
|
||||
node_test_failures_30d: number;
|
||||
alerts_fired_30d: number;
|
||||
|
|
|
|||
|
|
@ -83,6 +83,13 @@ const buildTelemetryPreviewPayload = (
|
|||
node_test_attempts_30d: 0,
|
||||
node_test_failures_30d: 0,
|
||||
update_last_failure_category: undefined,
|
||||
service_health_observed: true,
|
||||
service_health_healthy: true,
|
||||
service_health_failure_category: undefined,
|
||||
service_health_cohort: 'same_version',
|
||||
service_health_previous_version: '6.3.0',
|
||||
service_health_previous_observed: true,
|
||||
service_health_previous_healthy: true,
|
||||
alerts_fired_30d: 0,
|
||||
alerts_acknowledged_30d: 0,
|
||||
alerts_resolved_30d: 0,
|
||||
|
|
|
|||
172
internal/telemetry/service_health.go
Normal file
172
internal/telemetry/service_health.go
Normal file
|
|
@ -0,0 +1,172 @@
|
|||
package telemetry
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strings"
|
||||
"sync"
|
||||
|
||||
"github.com/rs/zerolog/log"
|
||||
)
|
||||
|
||||
type serviceHealthRecord struct {
|
||||
SchemaVersion int `json:"schema_version"`
|
||||
CurrentVersion string `json:"current_version"`
|
||||
CurrentObserved bool `json:"current_observed"`
|
||||
CurrentHealthy bool `json:"current_healthy"`
|
||||
CurrentFailureCategory string `json:"current_failure_category,omitempty"`
|
||||
PreviousVersion string `json:"previous_version,omitempty"`
|
||||
PreviousObserved bool `json:"previous_observed"`
|
||||
PreviousHealthy bool `json:"previous_healthy"`
|
||||
PreviousFailureCategory string `json:"previous_failure_category,omitempty"`
|
||||
}
|
||||
|
||||
var serviceHealthMu sync.Mutex
|
||||
|
||||
var storedServiceVersionPattern = regexp.MustCompile(`^\d+\.\d+\.\d+(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$`)
|
||||
|
||||
func applyServiceHealth(ping *Ping, dataDir string, observe func() ServiceHealthObservation) {
|
||||
if ping == nil || observe == nil {
|
||||
return
|
||||
}
|
||||
|
||||
observation := normalizeServiceHealthObservation(observe())
|
||||
serviceHealthMu.Lock()
|
||||
defer serviceHealthMu.Unlock()
|
||||
|
||||
record := readServiceHealthRecord(dataDir)
|
||||
switch {
|
||||
case record.CurrentVersion == "":
|
||||
ping.ServiceHealthCohort = ServiceHealthCohortFirstObservation
|
||||
case record.CurrentVersion != ping.Version:
|
||||
ping.ServiceHealthCohort = ServiceHealthCohortVersionChange
|
||||
record.PreviousVersion = record.CurrentVersion
|
||||
record.PreviousObserved = record.CurrentObserved
|
||||
record.PreviousHealthy = record.CurrentHealthy
|
||||
record.PreviousFailureCategory = record.CurrentFailureCategory
|
||||
default:
|
||||
ping.ServiceHealthCohort = ServiceHealthCohortSameVersion
|
||||
}
|
||||
|
||||
record.SchemaVersion = 1
|
||||
record.CurrentVersion = ping.Version
|
||||
record.CurrentObserved = observation.Observed
|
||||
record.CurrentHealthy = observation.Healthy
|
||||
record.CurrentFailureCategory = observation.FailureCategory
|
||||
|
||||
ping.ServiceHealthObserved = observation.Observed
|
||||
ping.ServiceHealthHealthy = observation.Healthy
|
||||
ping.ServiceHealthFailureCategory = observation.FailureCategory
|
||||
ping.ServiceHealthPreviousVersion = record.PreviousVersion
|
||||
ping.ServiceHealthPreviousObserved = record.PreviousObserved
|
||||
ping.ServiceHealthPreviousHealthy = record.PreviousHealthy
|
||||
|
||||
if err := writeServiceHealthRecord(dataDir, record); err != nil {
|
||||
log.Debug().Err(err).Msg("Could not persist coarse telemetry service-health observation")
|
||||
}
|
||||
}
|
||||
|
||||
func normalizeServiceHealthObservation(observation ServiceHealthObservation) ServiceHealthObservation {
|
||||
if !observation.Observed {
|
||||
return ServiceHealthObservation{}
|
||||
}
|
||||
if observation.Healthy {
|
||||
return ServiceHealthObservation{Observed: true, Healthy: true}
|
||||
}
|
||||
observation.FailureCategory = canonicalServiceHealthFailureCategory(observation.FailureCategory)
|
||||
return observation
|
||||
}
|
||||
|
||||
func canonicalServiceHealthFailureCategory(value string) string {
|
||||
switch strings.ToLower(strings.TrimSpace(value)) {
|
||||
case ServiceHealthFailureListener,
|
||||
ServiceHealthFailureStartup,
|
||||
ServiceHealthFailureRuntime,
|
||||
ServiceHealthFailureAPIConnectivity,
|
||||
ServiceHealthFailureAPIStatus,
|
||||
ServiceHealthFailureUIStatus,
|
||||
ServiceHealthFailureFrontendAssets:
|
||||
return strings.ToLower(strings.TrimSpace(value))
|
||||
default:
|
||||
return ServiceHealthFailureUnknown
|
||||
}
|
||||
}
|
||||
|
||||
func readServiceHealthRecord(dataDir string) serviceHealthRecord {
|
||||
data, err := os.ReadFile(filepath.Join(dataDir, serviceHealthStateFile))
|
||||
if err != nil {
|
||||
return serviceHealthRecord{}
|
||||
}
|
||||
var record serviceHealthRecord
|
||||
if err := json.Unmarshal(data, &record); err != nil || record.SchemaVersion != 1 {
|
||||
return serviceHealthRecord{}
|
||||
}
|
||||
record.CurrentVersion = canonicalStoredServiceVersion(record.CurrentVersion)
|
||||
if record.CurrentVersion == "" {
|
||||
record.CurrentObserved = false
|
||||
record.CurrentHealthy = false
|
||||
record.CurrentFailureCategory = ""
|
||||
}
|
||||
record.PreviousVersion = canonicalStoredServiceVersion(record.PreviousVersion)
|
||||
if record.PreviousVersion == "" {
|
||||
record.PreviousObserved = false
|
||||
record.PreviousHealthy = false
|
||||
record.PreviousFailureCategory = ""
|
||||
}
|
||||
record.CurrentFailureCategory = storedServiceHealthFailureCategory(
|
||||
record.CurrentObserved,
|
||||
record.CurrentHealthy,
|
||||
record.CurrentFailureCategory,
|
||||
)
|
||||
record.PreviousFailureCategory = storedServiceHealthFailureCategory(
|
||||
record.PreviousObserved,
|
||||
record.PreviousHealthy,
|
||||
record.PreviousFailureCategory,
|
||||
)
|
||||
return record
|
||||
}
|
||||
|
||||
func canonicalStoredServiceVersion(value string) string {
|
||||
value = strings.TrimPrefix(strings.TrimSpace(value), "v")
|
||||
if len(value) > 64 || !storedServiceVersionPattern.MatchString(value) {
|
||||
return ""
|
||||
}
|
||||
return value
|
||||
}
|
||||
|
||||
func storedServiceHealthFailureCategory(observed, healthy bool, category string) string {
|
||||
if !observed || healthy {
|
||||
return ""
|
||||
}
|
||||
return canonicalServiceHealthFailureCategory(category)
|
||||
}
|
||||
|
||||
func writeServiceHealthRecord(dataDir string, record serviceHealthRecord) error {
|
||||
if err := os.MkdirAll(dataDir, 0o700); err != nil {
|
||||
return err
|
||||
}
|
||||
encoded, err := json.Marshal(record)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
tmp, err := os.CreateTemp(dataDir, ".telemetry_service_health-*")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
tmpPath := tmp.Name()
|
||||
defer os.Remove(tmpPath)
|
||||
if err := tmp.Chmod(0o600); err != nil {
|
||||
_ = tmp.Close()
|
||||
return err
|
||||
}
|
||||
if _, err := tmp.Write(append(encoded, '\n')); err != nil {
|
||||
_ = tmp.Close()
|
||||
return err
|
||||
}
|
||||
if err := tmp.Close(); err != nil {
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmpPath, filepath.Join(dataDir, serviceHealthStateFile))
|
||||
}
|
||||
167
internal/telemetry/service_health_test.go
Normal file
167
internal/telemetry/service_health_test.go
Normal file
|
|
@ -0,0 +1,167 @@
|
|||
package telemetry
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func TestSendServiceHealthEventSendsImmediateStartupFailure(t *testing.T) {
|
||||
var received Ping
|
||||
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if err := json.NewDecoder(r.Body).Decode(&received); err != nil {
|
||||
t.Errorf("decode telemetry ping: %v", err)
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}))
|
||||
defer server.Close()
|
||||
originalEndpoint := pingEndpoint
|
||||
pingEndpoint = server.URL
|
||||
defer func() { pingEndpoint = originalEndpoint }()
|
||||
|
||||
err := SendServiceHealthEvent(context.Background(), Config{
|
||||
Version: "6.5.0",
|
||||
DataDir: t.TempDir(),
|
||||
Enabled: true,
|
||||
}, "startup", ServiceHealthObservation{
|
||||
Observed: true,
|
||||
FailureCategory: ServiceHealthFailureListener,
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("send immediate service-health event: %v", err)
|
||||
}
|
||||
if received.Event != "startup" || !received.ServiceHealthObserved ||
|
||||
received.ServiceHealthHealthy || received.ServiceHealthFailureCategory != ServiceHealthFailureListener {
|
||||
t.Fatalf("immediate service-health ping = %#v", received)
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildPingAtServiceHealthVersionChangeCohort(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
now := time.Date(2026, 8, 29, 10, 0, 0, 0, time.UTC)
|
||||
observation := ServiceHealthObservation{Observed: true, Healthy: true}
|
||||
cfg := Config{
|
||||
Version: "6.4.0",
|
||||
DataDir: dir,
|
||||
GetServiceHealth: func() ServiceHealthObservation { return observation },
|
||||
}
|
||||
|
||||
first, err := buildPingAt(cfg, "startup", now)
|
||||
if err != nil {
|
||||
t.Fatalf("build first ping: %v", err)
|
||||
}
|
||||
if !first.ServiceHealthObserved || !first.ServiceHealthHealthy || first.ServiceHealthCohort != ServiceHealthCohortFirstObservation {
|
||||
t.Fatalf("first service-health observation = %#v", first)
|
||||
}
|
||||
|
||||
observation = ServiceHealthObservation{Observed: true, FailureCategory: ServiceHealthFailureFrontendAssets}
|
||||
cfg.Version = "6.5.0"
|
||||
upgraded, err := buildPingAt(cfg, "startup", now.Add(time.Hour))
|
||||
if err != nil {
|
||||
t.Fatalf("build upgraded ping: %v", err)
|
||||
}
|
||||
if upgraded.ServiceHealthHealthy || upgraded.ServiceHealthFailureCategory != ServiceHealthFailureFrontendAssets {
|
||||
t.Fatalf("upgraded service-health observation = %#v", upgraded)
|
||||
}
|
||||
if upgraded.ServiceHealthCohort != ServiceHealthCohortVersionChange ||
|
||||
upgraded.ServiceHealthPreviousVersion != "6.4.0" ||
|
||||
!upgraded.ServiceHealthPreviousObserved ||
|
||||
!upgraded.ServiceHealthPreviousHealthy {
|
||||
t.Fatalf("version-change cohort = %#v", upgraded)
|
||||
}
|
||||
|
||||
observation = ServiceHealthObservation{Observed: true, Healthy: true, FailureCategory: "must-not-survive"}
|
||||
recovered, err := buildPingAt(cfg, "heartbeat", now.Add(2*time.Hour))
|
||||
if err != nil {
|
||||
t.Fatalf("build recovered ping: %v", err)
|
||||
}
|
||||
if recovered.ServiceHealthCohort != ServiceHealthCohortSameVersion ||
|
||||
!recovered.ServiceHealthHealthy || recovered.ServiceHealthFailureCategory != "" ||
|
||||
recovered.ServiceHealthPreviousVersion != "6.4.0" ||
|
||||
!recovered.ServiceHealthPreviousHealthy {
|
||||
t.Fatalf("same-version recovery = %#v", recovered)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceHealthObservationRejectsFreeFormFailure(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
cfg := Config{
|
||||
Version: "6.5.0",
|
||||
DataDir: dir,
|
||||
GetServiceHealth: func() ServiceHealthObservation {
|
||||
return ServiceHealthObservation{
|
||||
Observed: true,
|
||||
FailureCategory: "GET http://10.0.0.4:7655/assets/private.js: connection refused",
|
||||
}
|
||||
},
|
||||
}
|
||||
ping, err := buildPingAt(cfg, "startup", time.Now().UTC())
|
||||
if err != nil {
|
||||
t.Fatalf("build ping: %v", err)
|
||||
}
|
||||
if ping.ServiceHealthFailureCategory != ServiceHealthFailureUnknown {
|
||||
t.Fatalf("failure category = %q, want unknown", ping.ServiceHealthFailureCategory)
|
||||
}
|
||||
for _, forbidden := range []string{"http", "10.0.0.4", "private.js", "connection refused"} {
|
||||
if strings.Contains(ping.ServiceHealthFailureCategory, forbidden) {
|
||||
t.Fatalf("failure category leaked %q: %q", forbidden, ping.ServiceHealthFailureCategory)
|
||||
}
|
||||
}
|
||||
|
||||
raw, err := os.ReadFile(filepath.Join(dir, serviceHealthStateFile))
|
||||
if err != nil {
|
||||
t.Fatalf("read service-health state: %v", err)
|
||||
}
|
||||
if strings.Contains(string(raw), "10.0.0.4") || strings.Contains(string(raw), "private.js") {
|
||||
t.Fatalf("service-health state leaked free-form detail: %s", raw)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceHealthStateFileIsOwnerOnly(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
cfg := Config{
|
||||
Version: "6.5.0",
|
||||
DataDir: dir,
|
||||
GetServiceHealth: func() ServiceHealthObservation {
|
||||
return ServiceHealthObservation{Observed: true, Healthy: true}
|
||||
},
|
||||
}
|
||||
if _, err := buildPingAt(cfg, "startup", time.Now().UTC()); err != nil {
|
||||
t.Fatalf("build ping: %v", err)
|
||||
}
|
||||
info, err := os.Stat(filepath.Join(dir, serviceHealthStateFile))
|
||||
if err != nil {
|
||||
t.Fatalf("stat service-health state: %v", err)
|
||||
}
|
||||
if got := info.Mode().Perm(); got != 0o600 {
|
||||
t.Fatalf("service-health state mode = %o, want 600", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceHealthStateRejectsTamperedPreviousVersion(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
state := `{"schema_version":1,"current_version":"https://customer.example/private","current_observed":true,"current_healthy":true}`
|
||||
if err := os.WriteFile(filepath.Join(dir, serviceHealthStateFile), []byte(state), 0o600); err != nil {
|
||||
t.Fatalf("write tampered state: %v", err)
|
||||
}
|
||||
cfg := Config{
|
||||
Version: "6.5.0",
|
||||
DataDir: dir,
|
||||
GetServiceHealth: func() ServiceHealthObservation {
|
||||
return ServiceHealthObservation{Observed: true, Healthy: true}
|
||||
},
|
||||
}
|
||||
ping, err := buildPingAt(cfg, "startup", time.Now().UTC())
|
||||
if err != nil {
|
||||
t.Fatalf("build ping: %v", err)
|
||||
}
|
||||
if ping.ServiceHealthPreviousVersion != "" || ping.ServiceHealthPreviousObserved || ping.ServiceHealthPreviousHealthy {
|
||||
t.Fatalf("tampered previous release escaped local boundary: %#v", ping)
|
||||
}
|
||||
}
|
||||
|
|
@ -18,6 +18,8 @@
|
|||
// - Known install age, highest activation stage, time to first monitored resource, and estate-size buckets
|
||||
// - Whether authentication is configured, number of configured connections, and whether monitoring is active
|
||||
// - Whether a core outcome was observed in the current aggregate windows
|
||||
// - Whether the local API, UI document, and referenced frontend assets are
|
||||
// being served, plus the immediately previous release observation
|
||||
//
|
||||
// Scale (counts only, no names):
|
||||
// - Number of PVE nodes, PBS instances, PMG instances
|
||||
|
|
@ -58,6 +60,7 @@
|
|||
// - No alert content, AI prompts, chat messages, command text, action output, or token values
|
||||
// - No action targets, resource IDs, finding IDs, approval actors, or approval reasons
|
||||
// - No names, email addresses, account identifiers, or other intentionally identifying personal content
|
||||
// - No listener addresses, self-check URLs, response bodies, asset names, or raw startup/self-check errors
|
||||
//
|
||||
// # How to disable
|
||||
//
|
||||
|
|
@ -123,6 +126,11 @@ const (
|
|||
// identifiers and is intentionally independent from the rotating install ID.
|
||||
lifecycleStateFile = ".telemetry_lifecycle"
|
||||
|
||||
// serviceHealthStateFile retains only the current and immediately previous
|
||||
// release observation. It lets a post-upgrade ping describe a before/after
|
||||
// cohort without exporting listener addresses, URLs, errors, or request data.
|
||||
serviceHealthStateFile = ".telemetry_service_health"
|
||||
|
||||
// installIDRotationWindow limits how long the same pseudonymous identifier
|
||||
// can be reused before it is rotated locally.
|
||||
installIDRotationWindow = 30 * 24 * time.Hour
|
||||
|
|
@ -168,7 +176,11 @@ const (
|
|||
// Schema v12 adds Patrol-origin action funnel counters so detection and
|
||||
// investigation activity can be separated from proposal, decision,
|
||||
// execution, and successful completion without exporting action identity.
|
||||
TelemetrySchemaVersion = 12
|
||||
// Schema v13 adds a local UI/API service observation and the immediately
|
||||
// previous release observation. This separates a process that can emit
|
||||
// telemetry from one that is actually serving its API and frontend assets,
|
||||
// while retaining only fixed categories and release identity.
|
||||
TelemetrySchemaVersion = 13
|
||||
)
|
||||
|
||||
type installIDRecord struct {
|
||||
|
|
@ -182,6 +194,31 @@ type lifecycleRecord struct {
|
|||
HighestObservedActivation string `json:"highest_observed_activation"`
|
||||
}
|
||||
|
||||
// ServiceHealthObservation is the bounded result of Pulse's local UI/API
|
||||
// self-check. FailureCategory must be one of the fixed ServiceHealthFailure*
|
||||
// values. Callers must never put an address, URL, error string, or response
|
||||
// content into this shape.
|
||||
type ServiceHealthObservation struct {
|
||||
Observed bool
|
||||
Healthy bool
|
||||
FailureCategory string
|
||||
}
|
||||
|
||||
const (
|
||||
ServiceHealthFailureListener = "listener"
|
||||
ServiceHealthFailureStartup = "startup"
|
||||
ServiceHealthFailureRuntime = "runtime"
|
||||
ServiceHealthFailureAPIConnectivity = "api_connectivity"
|
||||
ServiceHealthFailureAPIStatus = "api_status"
|
||||
ServiceHealthFailureUIStatus = "ui_status"
|
||||
ServiceHealthFailureFrontendAssets = "frontend_assets"
|
||||
ServiceHealthFailureUnknown = "unknown"
|
||||
|
||||
ServiceHealthCohortFirstObservation = "first_observation"
|
||||
ServiceHealthCohortSameVersion = "same_version"
|
||||
ServiceHealthCohortVersionChange = "version_change"
|
||||
)
|
||||
|
||||
// Ping is the payload sent to the telemetry endpoint.
|
||||
// Every field is documented here so users can audit exactly what leaves their server.
|
||||
type Ping struct {
|
||||
|
|
@ -271,6 +308,17 @@ type Ping struct {
|
|||
// Last coarse update failure category; never raw error text.
|
||||
UpdateLastFailureCategory string `json:"update_last_failure_category,omitempty"`
|
||||
|
||||
// Local release-service observation. The probe checks the local API, UI
|
||||
// document, and referenced frontend assets. Only booleans, a fixed failure
|
||||
// category, and normalized release versions leave the instance.
|
||||
ServiceHealthObserved bool `json:"service_health_observed"`
|
||||
ServiceHealthHealthy bool `json:"service_health_healthy"`
|
||||
ServiceHealthFailureCategory string `json:"service_health_failure_category,omitempty"`
|
||||
ServiceHealthCohort string `json:"service_health_cohort,omitempty"`
|
||||
ServiceHealthPreviousVersion string `json:"service_health_previous_version,omitempty"`
|
||||
ServiceHealthPreviousObserved bool `json:"service_health_previous_observed"`
|
||||
ServiceHealthPreviousHealthy bool `json:"service_health_previous_healthy"`
|
||||
|
||||
// Node connection test outcomes over the install-ID rotation window.
|
||||
// ConfiguredConnections counts only connections that were saved, so an
|
||||
// install that tried to reach a node and could not is indistinguishable
|
||||
|
|
@ -836,6 +884,7 @@ type Config struct {
|
|||
DeploymentMethod string
|
||||
Enabled bool // From cfg.TelemetryEnabled (system settings or env var)
|
||||
GetSnapshot SnapshotFunc
|
||||
GetServiceHealth func() ServiceHealthObservation
|
||||
}
|
||||
|
||||
// runner holds the state for the background heartbeat goroutine.
|
||||
|
|
@ -878,7 +927,7 @@ func Start(ctx context.Context, cfg Config) {
|
|||
|
||||
log.Info().
|
||||
Str("platform", platformName(cfg.IsDocker)).
|
||||
Msg("Outbound usage telemetry enabled — sends a rotating pseudonymous install ID, version identity, coarse lifecycle buckets, aggregate resource/outcome counts, feature flags, and content-free Patrol, Assistant, and capability-API usage counters")
|
||||
Msg("Outbound usage telemetry enabled: sends a rotating pseudonymous install ID, version identity, coarse lifecycle and local service-health buckets, aggregate resource/outcome counts, feature flags, and content-free Patrol, Assistant, and capability-API usage counters")
|
||||
|
||||
r.wg.Add(1)
|
||||
go func() {
|
||||
|
|
@ -928,6 +977,22 @@ func BuildPreview(cfg Config) (Ping, error) {
|
|||
return buildPingAt(cfg, "heartbeat", time.Now().UTC())
|
||||
}
|
||||
|
||||
// SendServiceHealthEvent sends one immediate bounded service-health ping. It
|
||||
// is used for listener/startup/runtime failures that would otherwise exit
|
||||
// before the normal delayed startup heartbeat. Disabled telemetry and mock
|
||||
// mode remain fully suppressed.
|
||||
func SendServiceHealthEvent(ctx context.Context, cfg Config, event string, observation ServiceHealthObservation) error {
|
||||
if !cfg.Enabled || mock.IsMockEnabled() {
|
||||
return nil
|
||||
}
|
||||
cfg.GetServiceHealth = func() ServiceHealthObservation { return observation }
|
||||
ping, err := buildPingAt(cfg, event, time.Now().UTC())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return send(ctx, ping)
|
||||
}
|
||||
|
||||
// ResetInstallID rotates the locally stored telemetry install ID immediately
|
||||
// and returns the new pseudonymous identifier.
|
||||
func ResetInstallID(dataDir string) (string, error) {
|
||||
|
|
@ -1462,6 +1527,7 @@ func buildPingAt(cfg Config, event string, now time.Time) (Ping, error) {
|
|||
ping.Event = event
|
||||
ping.SentAt = now.Format(time.RFC3339)
|
||||
applyLifecycle(&ping, cfg.DataDir, now)
|
||||
applyServiceHealth(&ping, cfg.DataDir, cfg.GetServiceHealth)
|
||||
return ping, nil
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -149,7 +149,7 @@ func bindRuntimeVersion(version string) {
|
|||
}
|
||||
|
||||
// Run starts the Pulse monitoring server.
|
||||
func Run(ctx context.Context, version string) error {
|
||||
func Run(ctx context.Context, version string) (runErr error) {
|
||||
bindRuntimeVersion(version)
|
||||
|
||||
// Initialize logger with baseline defaults for early startup logs
|
||||
|
|
@ -193,9 +193,37 @@ func Run(ctx context.Context, version string) error {
|
|||
// Initialize license public key for Pro feature validation
|
||||
pkglicensing.InitEmbeddedPublicKey()
|
||||
|
||||
// Resolve telemetry identity before binding so a listener/startup failure can
|
||||
// still report the same privacy-bounded service-health contract before exit.
|
||||
mtPersistence := config.NewMultiTenantPersistence(cfg.DataPath)
|
||||
baseDataDir := mtPersistence.BaseDataDir()
|
||||
isDocker := os.Getenv("PULSE_DOCKER") == "true"
|
||||
failureTelemetryCfg := telemetry.Config{
|
||||
Version: version,
|
||||
DataDir: baseDataDir,
|
||||
IsDocker: isDocker,
|
||||
Enabled: cfg.TelemetryEnabled,
|
||||
}
|
||||
failureCategory := telemetry.ServiceHealthFailureStartup
|
||||
defer func() {
|
||||
if runErr == nil {
|
||||
return
|
||||
}
|
||||
reportCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
|
||||
defer cancel()
|
||||
observation := telemetry.ServiceHealthObservation{
|
||||
Observed: true,
|
||||
FailureCategory: failureCategory,
|
||||
}
|
||||
if err := telemetry.SendServiceHealthEvent(reportCtx, failureTelemetryCfg, "startup", observation); err != nil {
|
||||
log.Debug().Err(err).Str("category", failureCategory).Msg("Could not send bounded service-health failure telemetry")
|
||||
}
|
||||
}()
|
||||
|
||||
mainAddr := fmt.Sprintf("%s:%d", cfg.BindAddress, cfg.FrontendPort)
|
||||
mainListener, err := net.Listen("tcp", mainAddr)
|
||||
if err != nil {
|
||||
failureCategory = telemetry.ServiceHealthFailureListener
|
||||
return fmt.Errorf("failed to bind UI/API server on %s: %w", mainAddr, err)
|
||||
}
|
||||
defer mainListener.Close()
|
||||
|
|
@ -208,16 +236,12 @@ func Run(ctx context.Context, version string) error {
|
|||
agentAddr := fmt.Sprintf("%s:%d", cfg.BindAddress, cfg.AgentIngestPort)
|
||||
agentListener, err = net.Listen("tcp", agentAddr)
|
||||
if err != nil {
|
||||
failureCategory = telemetry.ServiceHealthFailureListener
|
||||
return fmt.Errorf("failed to bind agent ingest server on %s: %w", agentAddr, err)
|
||||
}
|
||||
defer agentListener.Close()
|
||||
}
|
||||
|
||||
// Multi-tenant persistence is the canonical way to resolve the base data directory.
|
||||
// It uses cfg.DataPath, which already includes PULSE_DATA_DIR overrides.
|
||||
mtPersistence := config.NewMultiTenantPersistence(cfg.DataPath)
|
||||
baseDataDir := mtPersistence.BaseDataDir()
|
||||
|
||||
// Run multi-tenant data migration only when the feature is explicitly enabled.
|
||||
// This prevents any on-disk layout changes for default (single-tenant) users.
|
||||
if api.IsMultiTenantEnabled() {
|
||||
|
|
@ -484,13 +508,13 @@ func Run(ctx context.Context, version string) error {
|
|||
// Start pseudonymous telemetry (enabled by default; opt out via PULSE_TELEMETRY=false or Settings toggle).
|
||||
// Persistence is created once here (outside the closure) to avoid NewConfigPersistence's
|
||||
// fatal-on-error path running inside the telemetry goroutine.
|
||||
isDocker := os.Getenv("PULSE_DOCKER") == "true"
|
||||
telemetryPersistence := config.NewConfigPersistence(baseDataDir)
|
||||
telemetryCfg := telemetry.Config{
|
||||
Version: version,
|
||||
DataDir: baseDataDir,
|
||||
IsDocker: isDocker,
|
||||
Enabled: cfg.TelemetryEnabled,
|
||||
Version: version,
|
||||
DataDir: baseDataDir,
|
||||
IsDocker: isDocker,
|
||||
Enabled: cfg.TelemetryEnabled,
|
||||
GetServiceHealth: newServiceHealthProbe(mainListener, cfg.HTTPSEnabled && cfg.TLSCertFile != "" && cfg.TLSKeyFile != ""),
|
||||
GetSnapshot: func() telemetry.Snapshot {
|
||||
// Use the latest config (may have been swapped by a reload).
|
||||
currentCfg := cfg
|
||||
|
|
@ -626,6 +650,7 @@ func Run(ctx context.Context, version string) error {
|
|||
return snap
|
||||
},
|
||||
}
|
||||
failureTelemetryCfg = telemetryCfg
|
||||
telemetry.Start(ctx, telemetryCfg)
|
||||
defer telemetry.Stop()
|
||||
|
||||
|
|
@ -788,6 +813,7 @@ func Run(ctx context.Context, version string) error {
|
|||
serverErr <- err
|
||||
}
|
||||
}()
|
||||
failureCategory = telemetry.ServiceHealthFailureRuntime
|
||||
|
||||
if agentSrv != nil {
|
||||
go func() {
|
||||
|
|
@ -822,7 +848,6 @@ func Run(ctx context.Context, version string) error {
|
|||
defer signal.Stop(sigChan)
|
||||
defer signal.Stop(reloadChan)
|
||||
|
||||
var runErr error
|
||||
for {
|
||||
select {
|
||||
case err := <-serverErr:
|
||||
|
|
|
|||
162
pkg/server/service_health.go
Normal file
162
pkg/server/service_health.go
Normal file
|
|
@ -0,0 +1,162 @@
|
|||
package server
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/tls"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/rcourtman/pulse-go-rewrite/internal/telemetry"
|
||||
)
|
||||
|
||||
const (
|
||||
serviceHealthProbeTimeout = 5 * time.Second
|
||||
serviceHealthBodyLimit = 2 << 20
|
||||
serviceHealthAssetLimit = 32
|
||||
)
|
||||
|
||||
var frontendAssetReferencePattern = regexp.MustCompile(`(?i)(?:src|href)\s*=\s*["'](/assets/[^"'#?]+(?:\?[^"'#]*)?)["']`)
|
||||
|
||||
func newServiceHealthProbe(listener net.Listener, tlsEnabled bool) func() telemetry.ServiceHealthObservation {
|
||||
baseURL, ok := localServiceHealthBaseURL(listener, tlsEnabled)
|
||||
if !ok {
|
||||
return func() telemetry.ServiceHealthObservation {
|
||||
return telemetry.ServiceHealthObservation{
|
||||
Observed: true,
|
||||
FailureCategory: telemetry.ServiceHealthFailureAPIConnectivity,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
transport := &http.Transport{}
|
||||
if tlsEnabled {
|
||||
// The probe stays inside this process and connects only to the address
|
||||
// already bound by listener. Certificate trust is a client-facing concern,
|
||||
// while this probe verifies that Pulse can serve its own HTTPS handler.
|
||||
transport.TLSClientConfig = &tls.Config{InsecureSkipVerify: true} // #nosec G402
|
||||
}
|
||||
client := &http.Client{
|
||||
Transport: transport,
|
||||
Timeout: serviceHealthProbeTimeout,
|
||||
CheckRedirect: func(_ *http.Request, _ []*http.Request) error {
|
||||
return http.ErrUseLastResponse
|
||||
},
|
||||
}
|
||||
|
||||
return func() telemetry.ServiceHealthObservation {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), serviceHealthProbeTimeout)
|
||||
defer cancel()
|
||||
|
||||
apiBody, status, err := serviceHealthGET(ctx, client, baseURL+"/api/health")
|
||||
if err != nil {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureAPIConnectivity)
|
||||
}
|
||||
if status < http.StatusOK || status >= http.StatusMultipleChoices {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureAPIStatus)
|
||||
}
|
||||
var health struct {
|
||||
Status string `json:"status"`
|
||||
}
|
||||
if json.Unmarshal(apiBody, &health) != nil || health.Status != "healthy" {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureAPIStatus)
|
||||
}
|
||||
|
||||
indexBody, status, err := serviceHealthGET(ctx, client, baseURL+"/")
|
||||
if err != nil || status < http.StatusOK || status >= http.StatusMultipleChoices ||
|
||||
!strings.Contains(strings.ToLower(string(indexBody)), "<html") {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureUIStatus)
|
||||
}
|
||||
|
||||
assetPaths := frontendAssetPaths(indexBody)
|
||||
if len(assetPaths) == 0 {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureFrontendAssets)
|
||||
}
|
||||
for _, assetPath := range assetPaths {
|
||||
body, assetStatus, assetErr := serviceHealthGET(ctx, client, baseURL+assetPath)
|
||||
if assetErr != nil || assetStatus < http.StatusOK || assetStatus >= http.StatusMultipleChoices || len(body) == 0 {
|
||||
return unhealthyServiceObservation(telemetry.ServiceHealthFailureFrontendAssets)
|
||||
}
|
||||
}
|
||||
|
||||
return telemetry.ServiceHealthObservation{Observed: true, Healthy: true}
|
||||
}
|
||||
}
|
||||
|
||||
func localServiceHealthBaseURL(listener net.Listener, tlsEnabled bool) (string, bool) {
|
||||
if listener == nil {
|
||||
return "", false
|
||||
}
|
||||
tcpAddr, ok := listener.Addr().(*net.TCPAddr)
|
||||
if !ok || tcpAddr.Port <= 0 {
|
||||
return "", false
|
||||
}
|
||||
ip := tcpAddr.IP
|
||||
if ip == nil || ip.IsUnspecified() {
|
||||
if ip != nil && ip.To4() == nil {
|
||||
ip = net.IPv6loopback
|
||||
} else {
|
||||
ip = net.IPv4(127, 0, 0, 1)
|
||||
}
|
||||
}
|
||||
scheme := "http"
|
||||
if tlsEnabled {
|
||||
scheme = "https"
|
||||
}
|
||||
return fmt.Sprintf("%s://%s", scheme, net.JoinHostPort(ip.String(), fmt.Sprintf("%d", tcpAddr.Port))), true
|
||||
}
|
||||
|
||||
func frontendAssetPaths(index []byte) []string {
|
||||
matches := frontendAssetReferencePattern.FindAllSubmatch(index, serviceHealthAssetLimit)
|
||||
paths := make([]string, 0, len(matches))
|
||||
seen := make(map[string]struct{}, len(matches))
|
||||
for _, match := range matches {
|
||||
if len(match) != 2 {
|
||||
continue
|
||||
}
|
||||
parsed, err := url.Parse(string(match[1]))
|
||||
if err != nil || !strings.HasPrefix(parsed.Path, "/assets/") {
|
||||
continue
|
||||
}
|
||||
path := parsed.EscapedPath()
|
||||
if parsed.RawQuery != "" {
|
||||
path += "?" + parsed.RawQuery
|
||||
}
|
||||
if _, ok := seen[path]; ok {
|
||||
continue
|
||||
}
|
||||
seen[path] = struct{}{}
|
||||
paths = append(paths, path)
|
||||
}
|
||||
return paths
|
||||
}
|
||||
|
||||
func serviceHealthGET(ctx context.Context, client *http.Client, target string) ([]byte, int, error) {
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, target, nil)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
resp, err := client.Do(req)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
body, err := io.ReadAll(io.LimitReader(resp.Body, serviceHealthBodyLimit+1))
|
||||
if err != nil {
|
||||
return nil, resp.StatusCode, err
|
||||
}
|
||||
if len(body) > serviceHealthBodyLimit {
|
||||
return nil, resp.StatusCode, fmt.Errorf("service-health response exceeded bounded size")
|
||||
}
|
||||
return body, resp.StatusCode, nil
|
||||
}
|
||||
|
||||
func unhealthyServiceObservation(category string) telemetry.ServiceHealthObservation {
|
||||
return telemetry.ServiceHealthObservation{Observed: true, FailureCategory: category}
|
||||
}
|
||||
116
pkg/server/service_health_test.go
Normal file
116
pkg/server/service_health_test.go
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
package server
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"net"
|
||||
"net/http"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/rcourtman/pulse-go-rewrite/internal/telemetry"
|
||||
)
|
||||
|
||||
func TestServiceHealthProbeCoversAPIUIAndFrontendAssets(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
handler http.Handler
|
||||
wantHealthy bool
|
||||
wantCategory string
|
||||
}{
|
||||
{
|
||||
name: "healthy",
|
||||
handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
switch r.URL.Path {
|
||||
case "/api/health":
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
fmt.Fprint(w, `{"status":"healthy"}`)
|
||||
case "/":
|
||||
w.Header().Set("Content-Type", "text/html")
|
||||
fmt.Fprint(w, `<html><head><link rel="stylesheet" href="/assets/app.css"></head><body><script src="/assets/app.js"></script></body></html>`)
|
||||
case "/assets/app.css":
|
||||
fmt.Fprint(w, "body{}")
|
||||
case "/assets/app.js":
|
||||
fmt.Fprint(w, "console.log('ok')")
|
||||
default:
|
||||
http.NotFound(w, r)
|
||||
}
|
||||
}),
|
||||
wantHealthy: true,
|
||||
},
|
||||
{
|
||||
name: "api unhealthy",
|
||||
handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Path == "/api/health" {
|
||||
http.Error(w, `{"status":"unhealthy"}`, http.StatusServiceUnavailable)
|
||||
return
|
||||
}
|
||||
http.NotFound(w, r)
|
||||
}),
|
||||
wantCategory: telemetry.ServiceHealthFailureAPIStatus,
|
||||
},
|
||||
{
|
||||
name: "ui unavailable",
|
||||
handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Path == "/api/health" {
|
||||
fmt.Fprint(w, `{"status":"healthy"}`)
|
||||
return
|
||||
}
|
||||
http.NotFound(w, r)
|
||||
}),
|
||||
wantCategory: telemetry.ServiceHealthFailureUIStatus,
|
||||
},
|
||||
{
|
||||
name: "frontend asset unavailable",
|
||||
handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
switch r.URL.Path {
|
||||
case "/api/health":
|
||||
fmt.Fprint(w, `{"status":"healthy"}`)
|
||||
case "/":
|
||||
fmt.Fprint(w, `<html><body><script src="/assets/missing.js"></script></body></html>`)
|
||||
default:
|
||||
http.NotFound(w, r)
|
||||
}
|
||||
}),
|
||||
wantCategory: telemetry.ServiceHealthFailureFrontendAssets,
|
||||
},
|
||||
}
|
||||
|
||||
for _, test := range tests {
|
||||
t.Run(test.name, func(t *testing.T) {
|
||||
listener, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
server := &http.Server{Handler: test.handler, ReadHeaderTimeout: time.Second}
|
||||
done := make(chan struct{})
|
||||
go func() {
|
||||
defer close(done)
|
||||
_ = server.Serve(listener)
|
||||
}()
|
||||
t.Cleanup(func() {
|
||||
_ = server.Close()
|
||||
<-done
|
||||
})
|
||||
|
||||
got := newServiceHealthProbe(listener, false)()
|
||||
if !got.Observed || got.Healthy != test.wantHealthy || got.FailureCategory != test.wantCategory {
|
||||
t.Fatalf("service-health observation = %#v, want healthy=%v category=%q", got, test.wantHealthy, test.wantCategory)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceHealthProbeReportsConnectivityWithoutListener(t *testing.T) {
|
||||
got := newServiceHealthProbe(nil, false)()
|
||||
if !got.Observed || got.Healthy || got.FailureCategory != telemetry.ServiceHealthFailureAPIConnectivity {
|
||||
t.Fatalf("service-health observation = %#v", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFrontendAssetPathsStayLocalAndBounded(t *testing.T) {
|
||||
index := []byte(`<html><head><link href="https://cdn.example/app.css"><link href="/assets/app.css?v=1"></head><body><script src="/assets/app.js"></script><script src="/not-assets/private.js"></script></body></html>`)
|
||||
got := frontendAssetPaths(index)
|
||||
if len(got) != 2 || got[0] != "/assets/app.css?v=1" || got[1] != "/assets/app.js" {
|
||||
t.Fatalf("frontend asset paths = %#v", got)
|
||||
}
|
||||
}
|
||||
|
|
@ -114,6 +114,15 @@ USER_BASE_COUNT_FIELDS = (
|
|||
("notification_failures_rejected_7d", "Notification destination rejections (7d, schema v5+)"),
|
||||
("notification_failures_unknown_7d", "Notification unknown failures (7d, schema v5+)"),
|
||||
)
|
||||
SERVICE_HEALTH_ROW_FIELDS = (
|
||||
"service_health_observed",
|
||||
"service_health_healthy",
|
||||
"service_health_failure_category",
|
||||
"service_health_cohort",
|
||||
"service_health_previous_version",
|
||||
"service_health_previous_observed",
|
||||
"service_health_previous_healthy",
|
||||
)
|
||||
NOTIFICATION_FAILURE_COUNT_SIGNALS = (
|
||||
(
|
||||
"notification_attempt_failures_7d_schema_v2",
|
||||
|
|
@ -1034,6 +1043,7 @@ REPORT_ROW_COLUMNS = tuple(
|
|||
*(key for key, _ in USER_BASE_CATEGORY_FIELDS),
|
||||
*(key for key, _ in USER_BASE_BOOL_FIELDS),
|
||||
*(key for key, _ in USER_BASE_COUNT_FIELDS),
|
||||
*SERVICE_HEALTH_ROW_FIELDS,
|
||||
*(key for key, _ in PULSE_INTELLIGENCE_BOOL_FIELDS),
|
||||
*(key for key, _ in PULSE_INTELLIGENCE_COUNT_FIELDS),
|
||||
)
|
||||
|
|
@ -2961,6 +2971,82 @@ def summarize_target_version_coverage(
|
|||
}
|
||||
|
||||
|
||||
def summarize_target_release_service_health(
|
||||
latest_by_install: dict[str, dict[str, Any]],
|
||||
published_versions: set[str],
|
||||
target_version: str,
|
||||
*,
|
||||
now: datetime | None = None,
|
||||
window: timedelta = timedelta(days=7),
|
||||
) -> dict[str, Any]:
|
||||
current_time = now or datetime.now(timezone.utc)
|
||||
normalized_target = normalize_release_tag(target_version)
|
||||
target_rows = [
|
||||
row
|
||||
for row in latest_by_install.values()
|
||||
if current_time - parse_received_at(str(row["received_at"])) <= window
|
||||
and classify_row_version(row, published_versions).version == normalized_target
|
||||
]
|
||||
|
||||
failure_categories: Counter[str] = Counter()
|
||||
cohorts: Counter[str] = Counter()
|
||||
previous_versions: Counter[str] = Counter()
|
||||
transitions: Counter[str] = Counter()
|
||||
observed = 0
|
||||
healthy = 0
|
||||
comparable_version_changes = 0
|
||||
|
||||
for row in target_rows:
|
||||
if not parse_optional_bool(row.get("service_health_observed")):
|
||||
continue
|
||||
observed += 1
|
||||
current_healthy = parse_optional_bool(row.get("service_health_healthy"))
|
||||
if current_healthy:
|
||||
healthy += 1
|
||||
failure_categories["healthy"] += 1
|
||||
else:
|
||||
category = str(row.get("service_health_failure_category") or "unknown").strip() or "unknown"
|
||||
failure_categories[category] += 1
|
||||
|
||||
cohort = str(row.get("service_health_cohort") or "unknown").strip() or "unknown"
|
||||
cohorts[cohort] += 1
|
||||
if not parse_optional_bool(row.get("service_health_previous_observed")):
|
||||
continue
|
||||
|
||||
previous_version = normalize_release_tag(str(row.get("service_health_previous_version") or ""))
|
||||
if not previous_version or previous_version == normalized_target:
|
||||
continue
|
||||
comparable_version_changes += 1
|
||||
previous_versions[previous_version] += 1
|
||||
previous_healthy = parse_optional_bool(row.get("service_health_previous_healthy"))
|
||||
transition = (
|
||||
("healthy" if previous_healthy else "unhealthy")
|
||||
+ "_to_"
|
||||
+ ("healthy" if current_healthy else "unhealthy")
|
||||
)
|
||||
transitions[transition] += 1
|
||||
|
||||
return {
|
||||
"version": normalized_target,
|
||||
"window": "7d",
|
||||
"target_installs": len(target_rows),
|
||||
"observed_installs": observed,
|
||||
"unobserved_installs": len(target_rows) - observed,
|
||||
"healthy_installs": healthy,
|
||||
"unhealthy_installs": observed - healthy,
|
||||
"failure_categories": counter_entries(failure_categories, "category"),
|
||||
"cohorts": counter_entries(cohorts, "cohort"),
|
||||
"comparable_version_change_installs": comparable_version_changes,
|
||||
"previous_versions": counter_entries(previous_versions, "version"),
|
||||
"transitions": counter_entries(transitions, "transition"),
|
||||
"interpretation": (
|
||||
"These are direct latest per-install local service observations. "
|
||||
"Version-change transitions compare the retained immediately previous release observation "
|
||||
"with the target release and do not attribute rolling historical counters to the target."
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def summarize_rows(
|
||||
db_stats: dict[str, Any],
|
||||
rows: Iterable[dict[str, Any]],
|
||||
|
|
@ -3071,6 +3157,14 @@ def summarize_rows(
|
|||
)
|
||||
if target_version
|
||||
else None,
|
||||
"target_release_service_health_7d": summarize_target_release_service_health(
|
||||
latest_by_install,
|
||||
published_versions,
|
||||
target_version,
|
||||
now=current_time,
|
||||
)
|
||||
if target_version
|
||||
else None,
|
||||
"target_release_followup": target_release_followup,
|
||||
"active_latest": {
|
||||
"active_24h": latest_install_windows["24h"]["active_installs"],
|
||||
|
|
@ -3319,6 +3413,46 @@ def format_text(summary: dict[str, Any], repo: str, since_days: int) -> str:
|
|||
else:
|
||||
lines.append(" - none")
|
||||
|
||||
service_health = summary.get("target_release_service_health_7d")
|
||||
if service_health:
|
||||
lines.extend(
|
||||
[
|
||||
"",
|
||||
f"Target release local service health ({service_health['version']}, {service_health['window']}):",
|
||||
f"- target installs: {service_health['target_installs']}",
|
||||
f"- direct observations: {service_health['observed_installs']}",
|
||||
f"- no schema-v13 observation: {service_health['unobserved_installs']}",
|
||||
f"- healthy: {service_health['healthy_installs']}",
|
||||
f"- unhealthy: {service_health['unhealthy_installs']}",
|
||||
"- latest direct result categories:",
|
||||
]
|
||||
)
|
||||
categories = service_health.get("failure_categories", [])
|
||||
if categories:
|
||||
lines.extend(
|
||||
f" - {entry['category']}: {entry['installs']} install(s)"
|
||||
for entry in categories
|
||||
)
|
||||
else:
|
||||
lines.append(" - none")
|
||||
lines.append(
|
||||
"- comparable version-change observations: "
|
||||
f"{service_health['comparable_version_change_installs']}"
|
||||
)
|
||||
transitions = service_health.get("transitions", [])
|
||||
if transitions:
|
||||
lines.extend(
|
||||
f" - {entry['transition']}: {entry['installs']} install(s)"
|
||||
for entry in transitions
|
||||
)
|
||||
else:
|
||||
lines.append(" - none")
|
||||
lines.append(
|
||||
"- interpretation: direct local API/UI/asset observations only; the before/after "
|
||||
"cohort uses the immediately previous release observation and never rolls historical "
|
||||
"usage or update counters into the target release"
|
||||
)
|
||||
|
||||
target_followup = summary.get("target_release_followup")
|
||||
if target_followup:
|
||||
lines.extend(
|
||||
|
|
|
|||
|
|
@ -188,6 +188,7 @@ class TelemetryAdoptionReportTest(unittest.TestCase):
|
|||
self.assertTrue(recent_count_fields <= projected)
|
||||
self.assertIn("alert_ai_enabled", projected)
|
||||
self.assertIn("update_last_failure_category", projected)
|
||||
self.assertTrue(set(report.SERVICE_HEALTH_ROW_FIELDS) <= projected)
|
||||
self.assertNotIn("business_estate", projected)
|
||||
|
||||
specs = {entry["field"]: entry for entry in report.telemetry_signal_specs()}
|
||||
|
|
@ -519,6 +520,72 @@ class TelemetryAdoptionReportTest(unittest.TestCase):
|
|||
self.assertLess(report.compare_semver_precedence("6.2.1", "6.3.0-rc.3"), 0)
|
||||
self.assertEqual(report.compare_semver_precedence("6.3.0+build.2", "6.3.0"), 0)
|
||||
|
||||
def test_target_release_service_health_uses_direct_version_change_observations(self) -> None:
|
||||
now = datetime(2026, 8, 29, 12, tzinfo=timezone.utc)
|
||||
rows = {
|
||||
"healthy-upgrade": {
|
||||
"install_id": "healthy-upgrade",
|
||||
"received_at": "2026-08-29 11:00:00",
|
||||
"version": "6.5.0",
|
||||
"service_health_observed": 1,
|
||||
"service_health_healthy": 1,
|
||||
"service_health_cohort": "version_change",
|
||||
"service_health_previous_version": "6.4.0",
|
||||
"service_health_previous_observed": 1,
|
||||
"service_health_previous_healthy": 1,
|
||||
},
|
||||
"broken-upgrade": {
|
||||
"install_id": "broken-upgrade",
|
||||
"received_at": "2026-08-29 10:00:00",
|
||||
"version": "6.5.0",
|
||||
"service_health_observed": 1,
|
||||
"service_health_healthy": 0,
|
||||
"service_health_failure_category": "frontend_assets",
|
||||
"service_health_cohort": "version_change",
|
||||
"service_health_previous_version": "6.4.0",
|
||||
"service_health_previous_observed": 1,
|
||||
"service_health_previous_healthy": 1,
|
||||
# Rolling counters are deliberately irrelevant to this summary.
|
||||
"update_successes_30d": 99,
|
||||
},
|
||||
"legacy-target": {
|
||||
"install_id": "legacy-target",
|
||||
"received_at": "2026-08-29 09:00:00",
|
||||
"version": "6.5.0",
|
||||
"service_health_observed": 0,
|
||||
},
|
||||
"other-release": {
|
||||
"install_id": "other-release",
|
||||
"received_at": "2026-08-29 11:30:00",
|
||||
"version": "6.4.0",
|
||||
"service_health_observed": 1,
|
||||
"service_health_healthy": 0,
|
||||
"service_health_failure_category": "listener",
|
||||
},
|
||||
}
|
||||
|
||||
summary = report.summarize_target_release_service_health(
|
||||
rows,
|
||||
{"6.4.0", "6.5.0"},
|
||||
"6.5.0",
|
||||
now=now,
|
||||
)
|
||||
|
||||
self.assertEqual(summary["target_installs"], 3)
|
||||
self.assertEqual(summary["observed_installs"], 2)
|
||||
self.assertEqual(summary["unobserved_installs"], 1)
|
||||
self.assertEqual(summary["healthy_installs"], 1)
|
||||
self.assertEqual(summary["unhealthy_installs"], 1)
|
||||
self.assertEqual(summary["comparable_version_change_installs"], 2)
|
||||
self.assertEqual(
|
||||
{entry["transition"]: entry["installs"] for entry in summary["transitions"]},
|
||||
{"healthy_to_healthy": 1, "healthy_to_unhealthy": 1},
|
||||
)
|
||||
self.assertEqual(
|
||||
{entry["category"]: entry["installs"] for entry in summary["failure_categories"]},
|
||||
{"healthy": 1, "frontend_assets": 1},
|
||||
)
|
||||
|
||||
def test_target_release_followup_excludes_first_heartbeat_baselines_and_flags_rollbacks(self) -> None:
|
||||
now = datetime(2026, 8, 19, 12, tzinfo=timezone.utc)
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue