T7: metrics to central VictoriaMetrics host
closedParent: bug 75966cc (cell architecture tracking). Monitoring lives on a dedicated new host (metrics + log aggregation + UI); cells are scrape targets.
Goal
One VictoriaMetrics dashboard showing all cells, with alerts that page on cell degradation.
Scope (decompose as 7a/7b/7c)
- 7a sovrn instrumentation: promote
prometheus/client_golang(currently indirect ingo.mod) to direct; request latency/counts, verifier sweeps, store errors, healthz-check gauges;/metricsendpoint (localhost-bound, scraped via exporter/direct). - 7b dependency inventory: Stalwart
/metrics/prometheus(OSS: queue depth, auth failures, cert expiry per docs/05 T7); ZDS exposes no Prometheus endpoint found — verify, else process-level + probe-derived metrics (healthz check results, blob op counts). - 7c VM host role: VictoriaMetrics (+ logs), scrape configs per cell + relay, dashboards, alert pack (target-down, queue-depth, backup-freshness, cert-expiry).
Acceptance
- Staged failure on a cell fires the matching alert; dashboards cover every cell + relay from one UI.
5 Comments
Metrics Box Implementation Plan (bug 6cae3a0)
Goal: Provision
infra.sovrn.atvia Ansible as the central metrics+logs box (VictoriaMetrics + VictoriaLogs + vmalert + Alertmanager behind Caddy), wire every cell to push metrics (vmagent) and forward logs (Vector from journald), and ship vmui dashboards + alerts covering disk space, queues, and health checks, with logs decomposable by cell and service.Architecture: No inbound scrape ports on cells. Each cell runs
vmagent(:8429loopback) scraping local exporters andremoteWrite-pushing to central VM with acell=external label, plusVectortailing systemd journald and shipping to central VL withcell/servicestream fields. Central box runs single-node VictoriaMetrics (:8428loopback) + VictoriaLogs (:9428loopback) + vmalert (:8880) + Alertmanager (:9093), all fronted by Caddy on 443 with basic-auth (one fleet shared secret). Stalwart is unified from/var/log/stalwartfiles to journald so one shipper covers everything. UI is built-in vmui day-1; Grafana explicitly deferred.Tech Stack: VictoriaMetrics single-server, VictoriaLogs, vmagent, vmalert, Alertmanager, Vector, node_exporter, Caddy, systemd journald, Ansible, MetricsQL/LogsQL, Resend SMTP for alert mail.
Decisions locked with operator (2026-09-16): vmagent push (not central scrape) · Vector journald→VL (not journal-upload/vlagent) · Stalwart to journald (not file tailing) · vmui only first · host
infra.sovrn.at(ownmetrics:group, monitoring-only for now; future CI/CD runner must not share data dirs) · basic-auth shared secret for ingest · alerts via Resend SMTP to[email protected](host_vars-configurable) · sovrn/metricsendpoint assumed to exist; its instrumentation is a separate child issue. Auth posture (locked 2026-09-16): basic-auth lives in Caddy ONLY. Do NOT set native-httpAuth.username/passwordflags on VM/VL/vmalert/Alertmanager — localhost trust on the box is accepted (future runner access to loopback is not a concern). Operator reaches vmui athttps://infra.sovrn.at/vmui/via the browser’s native basic-auth dialog with the fleet credential.Task 1: Inventory + secrets + versions groundwork
Files: - Modify:
deployment/inventory/hosts.yml- Create:deployment/inventory/host_vars/infra.sovrn.at/vars.yml- Modify:deployment/inventory/group_vars/all/sovrn.yml(addsovrn_metrics_*non-secret defaults) - Modify:deployment/inventory/group_vars/all/vault-shared.yml(vault-encrypted; operator step, never plaintext) - Modify:Justfile(addprovision-metrics-secretrecipe mirroringprovision-shared-secrets)Expected: six version+checksum pairs recorded;
ansible-playbook --checkstill passes (no role changes yet).Expected: dry-run OK;
head -c 15 deployment/inventory/group_vars/all/vault-shared.yml | grep '^\$ANSIBLE_VAULT'.Task 2:
metrics_boxrole — VM + VL + vmalert + Alertmanager + Caddy on infra.sovrn.atFiles: - Create:
deployment/roles/metrics_box/tasks/main.yml- Create:deployment/roles/metrics_box/handlers/main.yml- Create:deployment/roles/metrics_box/templates/victoriametrics.service.j2- Create:deployment/roles/metrics_box/templates/victorialogs.service.j2- Create:deployment/roles/metrics_box/templates/vmalert.service.j2- Create:deployment/roles/metrics_box/templates/alertmanager.service.j2- Create:deployment/roles/metrics_box/templates/alertmanager.yml.j2- Create:deployment/roles/metrics_box/templates/rules-cell.yml.j2- Create:deployment/roles/metrics_box/templates/rules-logs.yml.j2- Create:deployment/roles/metrics_box/templates/Caddyfile-infra.j2- Modify:deployment/playbooks/site.yml(gate box role togroups['metrics'], skip cell roles there)Expected: check-mode clean; converge green;
systemctl is-active victoriametrics victorialogs vmalert alertmanager caddyallactive.Expected:
vm_series present; unauthenticated write/query → 401; vmui loads athttps://infra.sovrn.at/vmui/.Task 3:
telemetry_edgerole — vmagent + Vector + node_exporter on every cellFiles: - Create:
deployment/roles/telemetry_edge/tasks/main.yml- Create:deployment/roles/telemetry_edge/handlers/main.yml- Create:deployment/roles/telemetry_edge/templates/scrape.yml.j2- Create:deployment/roles/telemetry_edge/templates/vmagent.service.j2- Create:deployment/roles/telemetry_edge/templates/vector.yaml.j2- Create:deployment/roles/telemetry_edge/templates/vector.service.j2- Create:deployment/roles/telemetry_edge/templates/node_exporter.service.j2- Modify:deployment/playbooks/site.yml(include edge role onmail_primary, tagtelemetry-edge)Expected:
up{cell="sovrn.at",service=~"host|stalwart|sovrnd"}present;sovrndtargetup==0until the child-issue/metricsendpoint lands (visible gap, tracked — does not block other targets).Expected: per-service counts;
zds@*user-unit lines present (proves root journald read covers user journals).Task 4: Stalwart journald unification (kill the file-tracer special case)
Files: - Modify:
deployment/roles/stalwart/files/bootstrap-stalwart.sh(tracer section) - Modify:deployment/inventory/group_vars/all/sovrn.yml(retire or repurposesovrn_stalwart_logdir) - Modify:deployment/roles/stalwart/tasks/install.yml(drop/var/log/stalwartcreation if present)Then replace the file tracer with the stdout equivalent (exact keys per checkout; stdout lines inherit
SYSLOG_IDENTIFIER=stalwartvia the systemd unit, and Vector maps_SYSTEMD_UNIT=stalwart.service→service="stalwart.service").Expected:
/var/log/stalwartabsent;journalctl -u stalwartshows live lines; VL query returns post-restart rows (no gap beyond restart window, proving the file-tail removal lost nothing).Task 5: Dashboards (vmui shareable queries) + staged-failure acceptance
Files: - Create:
docs/runbooks/metrics-box.md(vmui URLs, LogsQL/MetricsQL library, alert-response steps) - Modify:docs/deployment.md(infra box section: group, secrets, retention, UI URLs)Expected: each staged failure fires its matching alert to
[email protected]and resolves; results logged indocs/runbooks/metrics-box.mdwith timestamps.Task 6: Child issue + Grafana-deferral note
/metricsendpoint is created separately (agent runsbug agent new --parent 6cae3a0with the metric list below) — Task 3 scrape config already targets it, so this box plan stays unblocked.:3000on infra box with VM+VL datasources; no schema changes needed (labels are already Grafana-friendly).Self-review
cell=<inventory_hostname>,service=<job|SYSTEMD_UNIT>used identically in vmagent external_labels, Vector remap, VL stream fields, alert exprs, and dashboard queries.Implementation complete (12 commits, all spec+quality reviewed, final integration review: Ready to converge). Tasks 1-6 done: inventory+versions (VM 1.152.0, VL 1.52.0, Vector 0.58.0, node_exporter 1.12.1, AM 0.34.0, all checksummed), metrics_box role (VM+VL+dual vmalert+AM+Caddy basicauth-only on infra.sovrn.at), telemetry_edge role (vmagent push + Vector journald ship + node_exporter on cells), Stalwart->journald migration (live cells need manual JMAP x:Tracer migration, must NOT re-bootstrap), runbook docs/runbooks/metrics-box.md + staged-failure procedures, Grafana deferred. Child issues: dfb37b1 (/metrics endpoint), ba06600 (caddy_acme_ca undefined var, pre-existing). OPERATOR NEXT: 1) vault sovrn_metrics_password + bcrypt hash (just provision-metrics-secret) and confirm resend key; 2) DNS A for infra.sovrn.at; 3) triage sovrn.at SSH host-key change (do NOT blindly ssh-keygen -R); 4) just update infra.sovrn.at –tags metrics-box, then one canary cell –tags telemetry-edge; 5) run staged failures, log in runbook section 9. Static verification only — no live converge was possible from here.
Status (2026-10-06): the Ansible metrics box from 69c3a3e was superseded by the NixOS
metricsrole (dd3c64a); infra.sovrn.at was torn down and is rebuilt on the shared netcup box from~/projects/servers(plan Phase 8). Acceptance (a staged failure on a cell fires the matching alert) is still owed and needs a live cell: plan Phase 9. Cell-side telemetry is a6edca3; dfb37b1 (sovrnd /metrics) is needed before BackupStale/TlsExpiring can fire.Metrics box live on NixOS (2026-10-07, fleet Phase 8): infra.sovrn.at on the shared netcup box (infra.mymood.at), generation 15. LE cert, /healthz 200, UIs behind basic auth, test alert delivered via Resend. Fixed on the way (sovrn qtqttunr): Alertmanager now serves under /am, where Caddy proxies it (was a 404, as under Ansible). Remaining here: a staged failure firing the matching alert from a real cell (Phase 9).
Closing (user, 2026-10-08): acceptance met. The metrics box runs on NixOS (infra.sovrn.at); a staged failure on mx99 fired CellErrorBurst and reached [email protected]; vmui covers every cell, the relay and the shared box from one UI. Still open separately: dfb37b1 (sovrnd exports no sovrn_backup_marker_mtime_seconds / sovrn_tls_expiry_days yet, so BackupStale and TlsExpiring can’t fire).