Pulse/pkg/metrics
rcourtman 2c8cb8435d Make the time-major metrics index the unique identity index
Metric identity lived in idx_metrics_lookup, ordered (type, id, metric,
tier, timestamp). Once a series retains history, that order gives every
series its own insertion point, so each poll's commit rewrites one index
leaf page per series through the WAL and again at checkpoint. On a
three-node homelab running v6.4.5-rc.2 that came to 330-607 KB/s of disk
writes, about 10.6 KB per ~60-byte sample and 27-50 GB a day (#1966).

idx_metrics_query_all already holds the same five columns time-major, so
one poll's samples for a resource share pages. It now carries uniqueness
and serves every read, and the metric-major lookup index is dropped. With
an hour of retained history for 480 series, 30 polls write 6,022 WAL
frames instead of 25,310 (4.2x fewer); the empty-table issue-1124 estate
drops from 36,308 frames to 20,223, and its ceiling tightens from 40,000
to 23,000.

The tier-reconciliation overlap probe now names the identity index too;
left to the planner it seeks on (tier, timestamp) alone once the
metric-major tree is gone, which TestRetainedQueryPlansUseIndexes caught.

Metric-filtered reads now scan a resource's other metrics in the window:
a single-metric 500-node query went from 45 to 59 us and the 24-hour
50-disk smart_temp batch from about 0.25 to 0.49 s. All-metric dashboard
reads are unchanged. Continuous write volume is the user-visible harm, so
the contract records that trade explicitly.

The migration reuses the crash-safe swap: databases whose identity is
already enforced by a unique index defer the rebuild to startup
maintenance and swap in one transaction, and a database with no unique
index migrates synchronously. Keeping the name idx_metrics_query_all
means a downgrade rebuilds only the lookup index and leaves no orphan.

Refs #1966
2026-09-24 08:45:38 +01:00
..
alert_metrics.go
alert_metrics_test.go
availability_history.go Add availability history and fleet view 2026-08-30 15:38:34 +01:00
availability_history_test.go Add availability history and fleet view 2026-08-30 15:38:34 +01:00
directory_owner_other.go Protect shared metrics database directories 2026-08-11 15:57:17 +01:00
directory_owner_unix.go Protect shared metrics database directories 2026-08-11 15:57:17 +01:00
docker_observation_contract.go Preserve observed Docker storage evidence 2026-09-06 12:45:40 +01:00
identity_index_test.go Make the time-major metrics index the unique identity index 2026-09-24 08:45:38 +01:00
issue1124_write_amplification_test.go Make the time-major metrics index the unique identity index 2026-09-24 08:45:38 +01:00
norace_test.go
race_test.go
store.go Make the time-major metrics index the unique identity index 2026-09-24 08:45:38 +01:00
store_additional_test.go fix(metrics): survive startup write lock and bound allocations 2026-09-19 20:18:01 +01:00
store_bench_test.go
store_docker_observation_contract_test.go Preserve observed Docker storage evidence 2026-09-06 12:45:40 +01:00
store_downsample_test.go Prepare v6.0.0 release candidate 2026-06-04 14:07:14 +01:00
store_large_seed_test.go test(metrics): isolate large summary seed persistence across reopen 2026-09-06 15:59:08 +01:00
store_loadtest_test.go Bound metrics rollup working set 2026-07-03 20:35:02 +01:00
store_query_plan_test.go Make the time-major metrics index the unique identity index 2026-09-24 08:45:38 +01:00
store_series_coverage_test.go fix(mock): backfill the metrics store in mock mode so reports get real history 2026-06-10 21:25:02 +01:00
store_slo_test.go test(slo): only enforce latency budgets on hosted runners 2026-09-19 17:26:25 +01:00
store_test.go chore(metrics): move store regression proof into accepted proof file 2026-09-19 19:20:30 +01:00
store_tier_coverage_test.go perf(metrics): reuse retained query plans and output series 2026-09-06 05:36:19 +01:00
store_write_backpressure_test.go Split bounded pipeline metrics writes from synchronous batch writes 2026-08-11 14:42:42 +01:00