Telemetry edge on NixOS: vmagent, Vector, node_exporter
closedPort roles/telemetry_edge for cells and the relay:
- vmagent: scrape node_exporter, Stalwart /metrics/prometheus and sovrnd; push to infra
- Vector: journald to VictoriaLogs
- node_exporter
Use nixpkgs modules (services.vmagent, services.vector, services.prometheus.exporters.node). Scrape config and Vector config move from Jinja to Nix attrsets. Push credentials come from keys.
Done when infra shows metrics and logs from a NixOS cell, with labels matching the existing alert rules.
5 Comments
Decisions (2026-10-06, user): - The push endpoint is an option. Cells push to https://infra.sovrn.at (Caddy, basic auth). An edge on the same box as the metrics stack pushes straight to the loopback ingest (VictoriaMetrics 127.0.0.1:8428, VictoriaLogs 127.0.0.1:9428), no auth. That’s the shared netcup box, which runs the metrics stack, the relay and moods. - The
celllabel is the cell’s hostname on cells and the project name for everything else (moods,sovrn,servers). On the shared box, a unit/job → project map assigns it per service. - Export the module for the fleet (nixosModules.telemetryEdge).Done in NixOS (2026-10-07), pending live acceptance in Phases 8⁄9: - tnqrnrtk: nix/modules/telemetry-edge.nix (sovrn.telemetry.edge; nixosModules.telemetryEdge). node_exporter, vmagent, Vector on the nixpkgs modules, DynamicUser (password via LoadCredential, journals via systemd-journal group). Cells push to https:// with basic auth (metrics-password, declared by the edge); the metrics role’s own edge uses the loopback ingest without auth.
- Labels cell/service (+ level on logs) as before; set per scrape job. On a shared host: sovrnCell for sovrn’s units/jobs, unitCells (unit prefix -> cell) for other projects, cell for the rest. Log service prefers _SYSTEMD_USER_UNIT (ZDS instances named).
- checks.telemetry passes (cell via Caddy+auth, shared-box labelling); metrics, relay, cell, stalwart, stalwart-plan pass.
- nnsusooo: Caddy ACME email [email protected] (fleet.acme.caddyEmail); Stalwart’s ACME contact unchanged.
- Fleet (servers qwokvoqn): sovrn-metrics/sovrn-relay hosts label cell=servers, sovrn, moods.
Remaining for the done-condition: metrics and logs from a real NixOS cell (Phase 9, mx99).
Labels revised (user, 2026-10-07), folded into tnqrnrtk: every series and log line also carries host (VL stream field). cell = owner (alert rules group by it), host = where it runs; both the machine’s FQDN by default, so cells are distinct by hostname. On a shared host the relay labels its Stalwart and route-sync cell=host= (mxb.eu.sovrn.at), the metrics stack cell=host=infra.sovrn.at; other projects’ units by prefix (units.). sovrnCell/unitCells are gone. checks.relay now runs the metrics role beside the relay and asserts the relay’s Stalwart logs/metrics arrive as cell=host=. telemetry, relay, metrics, cell pass.
Edge live on the shared box (2026-10-07): host metrics cell=servers host=infra.mymood.at, moods units cell=moods, metrics stack cell=host=infra.sovrn.at, relay Stalwart cell=host=mxb.eu.sovrn.at (logs and up{job=stalwart}). Done-condition still needs a real NixOS cell (Phase 9).
Closing (user, 2026-10-08): done-condition met: infra.sovrn.at shows metrics and logs from mx99 (NixOS cell) with the labels the alert rules use; a staged CellErrorBurst fired and was emailed.