7a: sovrn /metrics endpoint (Prometheus exposition for vmagent scrape)
open7a: sovrn /metrics endpoint (Prometheus exposition for vmagent scrape)
Parent: bug 6cae3a0 (metrics box). The telemetry_edge vmagent scrape config targets http://127.0.0.1:8081/metrics; this issue tracks creating that endpoint. Until it lands, the sovrnd target reports up==0 — expected and visible, not silent.
Scope
- Promote
prometheus/client_golang(currently indirect ingo.modvia indigo, v1.17.0) to a direct dependency; pin version ingo.mod. - Add
GET /metricson the sovrnd top mux (router.go, next toGET /healthzat line ~223). Localhost-bound like:8081already is (behind Caddy); no auth (same posture as/healthz, loopback-only + vmagent scrape). - Register default Go/process collectors (
promhttpdefault) plus the sovrn series below. All sovrn series prefixedsovrn_; per-cell decomposition comes from vmagent’scell=external label, so no hostname label in-app.
Recommended initial series (minimal, matched to the alert pack + dashboards)
| Series | Type | Labels | Source | Consumed by |
|---|---|---|---|---|
sovrn_http_requests_total |
Counter | method, route, status |
existing logging.Middleware (method/path/route/status/duration/DID) in internal/logging |
latency/counts dashboards |
sovrn_http_request_duration_seconds |
Histogram | method, route |
same middleware (observe duration) | p99 dashboards, SLO burn later |
sovrn_health_check_ok |
Gauge | check (one per healthz leg: stalwart-ready, stalwart-jmap, zds- |
healthz.go check results (1⁄0 per evaluation) |
health drill-down without parsing JSON |
sovrn_backup_marker_mtime_seconds |
Gauge | marker (litestream, rclone) |
stat /var/lib/sovrn/health/*.ok mtimes (config.go HealthConfig paths) |
BackupStale alert (time() - ... > 2700) |
sovrn_tls_expiry_days |
Gauge | hostname |
TLS-expiry leg for served hostnames (recent healthz addition) |
TlsExpiring alert (< 14) |
sovrn_disk_used_ratio |
Gauge | path |
disk leg (>85% check) |
CellDiskHigh alert side (node_exporter remains primary; this is the app’s own view) |
sovrn_store_errors_total |
Counter | op, db (sovrn.db, oauth.db) |
internal/store error returns |
store-error dashboard + burst alert |
sovrn_verifier_sweep_total |
Counter | result (ok, fail) |
verifier sweep loop | sweep activity / stall detection (rate(...) == 0 over 2 intervals) |
sovrn_verifier_last_success_seconds |
Gauge | — | same loop (unixtime of last ok sweep) | stall alert without rate math |
Explicitly NOT in v1: per-DID/user/tenant labels (cardinality), per-PDS-instance blob op counts (ZDS exposes no Prometheus endpoint — covered by process-level node metrics + healthz-derived gauges per parent-bug 7b), trace exemplars.
Acceptance
curl 127.0.0.1:8081/metricson a converged cell exposessovrn_*+go_*/process_*series.- vmagent
up{job="sovrnd"}flips to 1 for that cell in the central VM. - New histogram has ≤10 buckets (default
prometheus.DefBucketsfine); no label with >50 values observed in smoke (check via/api/v1/status/tsdbcardinality or vmui cardinality explorer).
Notes for implementer
- Middleware already logs method/path/route/status/duration — attach histogram/counter there, using the matched route pattern (not raw path) as
routeto bound cardinality. - Healthz legs run on demand per request; gauges should be set as a side effect of each
/healthzevaluation AND/OR refreshed on a short ticker so/metricsstays fresh between healthz polls (vmagent scrapes every 15s; healthz polling cadence from monitoring should be ≥60s to avoid doubling Stalwart/JMAP probe load). - Keep
GET /healthzJSON body byte-identical (existing tests inhealthz_test.gopin it); only add gauge side-effects.
1 Comment
Status (2026-10-08, live on mx99 under NixOS): still not implemented. GET /metrics on sovrnd answers 303 to /login (the UI catch-all), so the telemetry edge’s sovrnd scrape job shows up=1 with scrape_samples_scraped=0: a misleading green. Consequences: BackupStale (sovrn_backup_marker_mtime_seconds) and TlsExpiring (sovrn_tls_expiry_days) can never fire, and there are no sovrnd request metrics. The design above still holds; the scrape side is already in place (nix/modules/telemetry-edge.nix job sovrnd -> 127.0.0.1:8081/metrics, cell label from the edge). Until it lands, consider follow_redirects: false on that scrape job so the gap shows as up=0.