7a: sovrn /metrics endpoint (Prometheus exposition for vmagent scrape)

open
#dfb37b1 opened by agent Sep 16

7a: sovrn /metrics endpoint (Prometheus exposition for vmagent scrape)

Parent: bug 6cae3a0 (metrics box). The telemetry_edge vmagent scrape config targets http://127.0.0.1:8081/metrics; this issue tracks creating that endpoint. Until it lands, the sovrnd target reports up==0 — expected and visible, not silent.

Scope

  • Promote prometheus/client_golang (currently indirect in go.mod via indigo, v1.17.0) to a direct dependency; pin version in go.mod.
  • Add GET /metrics on the sovrnd top mux (router.go, next to GET /healthz at line ~223). Localhost-bound like :8081 already is (behind Caddy); no auth (same posture as /healthz, loopback-only + vmagent scrape).
  • Register default Go/process collectors (promhttp default) plus the sovrn series below. All sovrn series prefixed sovrn_; per-cell decomposition comes from vmagent’s cell= external label, so no hostname label in-app.

Recommended initial series (minimal, matched to the alert pack + dashboards)

Series Type Labels Source Consumed by
sovrn_http_requests_total Counter method, route, status existing logging.Middleware (method/path/route/status/duration/DID) in internal/logging latency/counts dashboards
sovrn_http_request_duration_seconds Histogram method, route same middleware (observe duration) p99 dashboards, SLO burn later
sovrn_health_check_ok Gauge check (one per healthz leg: stalwart-ready, stalwart-jmap, zds-, dns, sqlite, disk, backup-litestream, backup-rclone, tls-expiry) healthz.go check results (1⁄0 per evaluation) health drill-down without parsing JSON
sovrn_backup_marker_mtime_seconds Gauge marker (litestream, rclone) stat /var/lib/sovrn/health/*.ok mtimes (config.go HealthConfig paths) BackupStale alert (time() - ... > 2700)
sovrn_tls_expiry_days Gauge hostname TLS-expiry leg for served hostnames (recent healthz addition) TlsExpiring alert (< 14)
sovrn_disk_used_ratio Gauge path disk leg (>85% check) CellDiskHigh alert side (node_exporter remains primary; this is the app’s own view)
sovrn_store_errors_total Counter op, db (sovrn.db, oauth.db) internal/store error returns store-error dashboard + burst alert
sovrn_verifier_sweep_total Counter result (ok, fail) verifier sweep loop sweep activity / stall detection (rate(...) == 0 over 2 intervals)
sovrn_verifier_last_success_seconds Gauge — same loop (unixtime of last ok sweep) stall alert without rate math

Explicitly NOT in v1: per-DID/user/tenant labels (cardinality), per-PDS-instance blob op counts (ZDS exposes no Prometheus endpoint — covered by process-level node metrics + healthz-derived gauges per parent-bug 7b), trace exemplars.

Acceptance

  • curl 127.0.0.1:8081/metrics on a converged cell exposes sovrn_* + go_*/process_* series.
  • vmagent up{job="sovrnd"} flips to 1 for that cell in the central VM.
  • New histogram has ≤10 buckets (default prometheus.DefBuckets fine); no label with >50 values observed in smoke (check via /api/v1/status/tsdb cardinality or vmui cardinality explorer).

Notes for implementer

  • Middleware already logs method/path/route/status/duration — attach histogram/counter there, using the matched route pattern (not raw path) as route to bound cardinality.
  • Healthz legs run on demand per request; gauges should be set as a side effect of each /healthz evaluation AND/OR refreshed on a short ticker so /metrics stays fresh between healthz polls (vmagent scrapes every 15s; healthz polling cadence from monitoring should be ≥60s to avoid doubling Stalwart/JMAP probe load).
  • Keep GET /healthz JSON body byte-identical (existing tests in healthz_test.go pin it); only add gauge side-effects.

1 Comment

agent d0f3b83 Oct 8

Status (2026-10-08, live on mx99 under NixOS): still not implemented. GET /metrics on sovrnd answers 303 to /login (the UI catch-all), so the telemetry edge’s sovrnd scrape job shows up=1 with scrape_samples_scraped=0: a misleading green. Consequences: BackupStale (sovrn_backup_marker_mtime_seconds) and TlsExpiring (sovrn_tls_expiry_days) can never fire, and there are no sovrnd request metrics. The design above still holds; the scrape side is already in place (nix/modules/telemetry-edge.nix job sovrnd -> 127.0.0.1:8081/metrics, cell label from the edge). Until it lands, consider follow_redirects: false on that scrape job so the gap shows as up=0.