T7: metrics to central VictoriaMetrics host

closed
#6cae3a0 opened by agent Sep 13

Parent: bug 75966cc (cell architecture tracking). Monitoring lives on a dedicated new host (metrics + log aggregation + UI); cells are scrape targets.

Goal

One VictoriaMetrics dashboard showing all cells, with alerts that page on cell degradation.

Scope (decompose as 7a/7b/7c)

  • 7a sovrn instrumentation: promote prometheus/client_golang (currently indirect in go.mod) to direct; request latency/counts, verifier sweeps, store errors, healthz-check gauges; /metrics endpoint (localhost-bound, scraped via exporter/direct).
  • 7b dependency inventory: Stalwart /metrics/prometheus (OSS: queue depth, auth failures, cert expiry per docs/05 T7); ZDS exposes no Prometheus endpoint found — verify, else process-level + probe-derived metrics (healthz check results, blob op counts).
  • 7c VM host role: VictoriaMetrics (+ logs), scrape configs per cell + relay, dashboards, alert pack (target-down, queue-depth, backup-freshness, cert-expiry).

Acceptance

  • Staged failure on a cell fires the matching alert; dashboards cover every cell + relay from one UI.

5 Comments

agent 6ec3a7e Sep 16

Metrics Box Implementation Plan (bug 6cae3a0)

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Provision infra.sovrn.at via Ansible as the central metrics+logs box (VictoriaMetrics + VictoriaLogs + vmalert + Alertmanager behind Caddy), wire every cell to push metrics (vmagent) and forward logs (Vector from journald), and ship vmui dashboards + alerts covering disk space, queues, and health checks, with logs decomposable by cell and service.

Architecture: No inbound scrape ports on cells. Each cell runs vmagent (:8429 loopback) scraping local exporters and remoteWrite-pushing to central VM with a cell= external label, plus Vector tailing systemd journald and shipping to central VL with cell/service stream fields. Central box runs single-node VictoriaMetrics (:8428 loopback) + VictoriaLogs (:9428 loopback) + vmalert (:8880) + Alertmanager (:9093), all fronted by Caddy on 443 with basic-auth (one fleet shared secret). Stalwart is unified from /var/log/stalwart files to journald so one shipper covers everything. UI is built-in vmui day-1; Grafana explicitly deferred.

Tech Stack: VictoriaMetrics single-server, VictoriaLogs, vmagent, vmalert, Alertmanager, Vector, node_exporter, Caddy, systemd journald, Ansible, MetricsQL/LogsQL, Resend SMTP for alert mail.

Decisions locked with operator (2026-09-16): vmagent push (not central scrape) · Vector journald→VL (not journal-upload/vlagent) · Stalwart to journald (not file tailing) · vmui only first · host infra.sovrn.at (own metrics: group, monitoring-only for now; future CI/CD runner must not share data dirs) · basic-auth shared secret for ingest · alerts via Resend SMTP to [email protected] (host_vars-configurable) · sovrn /metrics endpoint assumed to exist; its instrumentation is a separate child issue. Auth posture (locked 2026-09-16): basic-auth lives in Caddy ONLY. Do NOT set native -httpAuth.username/password flags on VM/VL/vmalert/Alertmanager — localhost trust on the box is accepted (future runner access to loopback is not a concern). Operator reaches vmui at https://infra.sovrn.at/vmui/ via the browser’s native basic-auth dialog with the fleet credential.


Task 1: Inventory + secrets + versions groundwork

Files: - Modify: deployment/inventory/hosts.yml - Create: deployment/inventory/host_vars/infra.sovrn.at/vars.yml - Modify: deployment/inventory/group_vars/all/sovrn.yml (add sovrn_metrics_* non-secret defaults) - Modify: deployment/inventory/group_vars/all/vault-shared.yml (vault-encrypted; operator step, never plaintext) - Modify: Justfile (add provision-metrics-secret recipe mirroring provision-shared-secrets)

  • [ ] Step 1: Add the metrics group to inventory
# deployment/inventory/hosts.yml — add alongside mail_primary
all:
  hosts:
    sovrn.at:
    infra.sovrn.at:
  children:
    mail_primary:
      hosts:
        sovrn.at:
    metrics:
      hosts:
        infra.sovrn.at:
  • [ ] Step 2: Create host vars for infra.sovrn.at
# deployment/inventory/host_vars/infra.sovrn.at/vars.yml
# Monitoring-only box (future CI/CD runner must use separate dirs/users).
sovrn_alert_recipient: "[email protected]"
sovrn_alert_from: "[email protected]"
sovrn_metrics_domain: "infra.sovrn.at"
sovrn_metrics_vm_retention: "90d"
sovrn_metrics_vl_retention: "30d"
sovrn_metrics_vl_max_disk: "80GiB"
  • [ ] Step 3: Add non-secret metrics defaults to group vars
# deployment/inventory/group_vars/all/sovrn.yml — append
# --- T7 metrics (bug 6cae3a0): non-secret defaults; secrets ONLY in vault-shared.yml
sovrn_metrics_vm_version: "<resolved in Step 4>"
sovrn_metrics_vl_version: "<resolved in Step 4>"
sovrn_metrics_vmagent_version: "<same as VM release>"
sovrn_metrics_vmalert_version: "<same as VM release>"
sovrn_metrics_vector_version: "<resolved in Step 4>"
sovrn_metrics_node_exporter_version: "<resolved in Step 4>"
sovrn_metrics_vm_port: 8428
sovrn_metrics_vl_port: 9428
sovrn_metrics_vmagent_port: 8429
sovrn_metrics_vmalert_port: 8880
sovrn_metrics_alertmanager_port: 9093
sovrn_metrics_node_port: 9100
sovrn_metrics_vm_datadir: "/var/lib/victoriametrics"
sovrn_metrics_vl_datadir: "/var/lib/victorialogs"
sovrn_metrics_ingest_user: "cell"
  • [ ] Step 4: Resolve pinned versions + checksums (no floating tags)
# On controller: pick current stable releases, record version + sha256 in sovrn.yml
curl -fsSL https://api.github.com/VictoriaMetrics/VictoriaMetrics/releases/latest | grep '"tag_name"'
curl -fsSL https://api.github.com/VictoriaMetrics/VictoriaMetrics/releases/latest | grep -o 'victoria-metrics-linux-amd64-.*tar.gz'
curl -fsSL https://api.github.com/vectordotdev/vector/releases/latest | grep '"tag_name"'
curl -fsSL https://api.github.com/prometheus/node_exporter/releases/latest | grep '"tag_name"'
# Download each asset, verify: sha256sum <file>, paste `sha256:<hex>` next to version vars

Expected: six version+checksum pairs recorded; ansible-playbook --check still passes (no role changes yet).

  • [ ] Step 5: Provision the shared ingest secret (operator step, documented not executed with real secret)
# Operator runs OUTSIDE the agent — never paste real credentials in chat:
#   secrets decrypt metrics_ingest_password   # or: openssl rand -base64 32
# then add to group_vars/all/vault-shared.yml:
#   sovrn_metrics_password: "<shared ingest + UI password>"
# and re-encrypt: ansible-vault encrypt deployment/inventory/group_vars/all/vault-shared.yml
# Cells receive it as /etc/vm/ingest-password (0600); infra box hashes it into Caddy basicauth.
just provision-shared-secrets --dry-run

Expected: dry-run OK; head -c 15 deployment/inventory/group_vars/all/vault-shared.yml | grep '^\$ANSIBLE_VAULT'.

  • [ ] Step 6: Commit
jj commit -m "feat(metrics): inventory group, host vars, version pins for infra box" deployment/inventory/hosts.yml deployment/inventory/host_vars/infra.sovrn.at/vars.yml deployment/inventory/group_vars/all/sovrn.yml Justfile

Task 2: metrics_box role — VM + VL + vmalert + Alertmanager + Caddy on infra.sovrn.at

Files: - Create: deployment/roles/metrics_box/tasks/main.yml - Create: deployment/roles/metrics_box/handlers/main.yml - Create: deployment/roles/metrics_box/templates/victoriametrics.service.j2 - Create: deployment/roles/metrics_box/templates/victorialogs.service.j2 - Create: deployment/roles/metrics_box/templates/vmalert.service.j2 - Create: deployment/roles/metrics_box/templates/alertmanager.service.j2 - Create: deployment/roles/metrics_box/templates/alertmanager.yml.j2 - Create: deployment/roles/metrics_box/templates/rules-cell.yml.j2 - Create: deployment/roles/metrics_box/templates/rules-logs.yml.j2 - Create: deployment/roles/metrics_box/templates/Caddyfile-infra.j2 - Modify: deployment/playbooks/site.yml (gate box role to groups['metrics'], skip cell roles there)

  • [ ] Step 1: Write role tasks (loopback-only daemons, Caddy terminates TLS)
# deployment/roles/metrics_box/tasks/main.yml (sketch — full file in implementation)
- name: Create metrics data dirs
  ansible.builtin.file: {path: "{{ item }}", state: directory, owner: sovrn, group: sovrn, mode: "0750"}
  loop: ["{{ sovrn_metrics_vm_datadir }}", "{{ sovrn_metrics_vl_datadir }}", /var/lib/vmagent-buffer, /etc/vmalert]
- name: Install pinned binaries (vm, vl, vmalert, alertmanager) with checksum verify
  ansible.builtin.get_url:
    url: "{{ item.url }}"
    dest: "/usr/local/sbin/{{ item.name }}"
    checksum: "{{ item.checksum }}"
    mode: "0755"
  loop: "{{ sovrn_metrics_binaries }}"
- name: Template systemd units + alertmanager.yml + vmalert rules
  ansible.builtin.template: {src: "{{ item.src }}", dest: "{{ item.dest }}", owner: root, group: root, mode: "0644"}
  loop:
    - {src: victoriametrics.service.j2, dest: /etc/systemd/system/victoriametrics.service}
    - {src: victorialogs.service.j2, dest: /etc/systemd/system/victorialogs.service}
    - {src: vmalert.service.j2, dest: /etc/systemd/system/vmalert.service}
    - {src: alertmanager.service.j2, dest: /etc/systemd/system/alertmanager.service}
  notify: Reload systemd
- name: Install Caddy vhost for infra box
  ansible.builtin.template: {src: Caddyfile-infra.j2, dest: /etc/caddy/Caddyfile, owner: root, group: root, mode: "0644"}
  notify: Reload caddy
- name: Open 80/443 only (no raw 8428/9428/8880/9093)
  ansible.builtin.ufw: {rule: allow, port: "{{ item }}", proto: tcp}
  loop: ["80", "443"]
# victoriametrics.service.j2 — ExecStart line
ExecStart=/usr/local/sbin/victoria-metrics -storageDataPath={{ sovrn_metrics_vm_datadir }} -retentionPeriod={{ sovrn_metrics_vm_retention }} -httpListenAddr=127.0.0.1:{{ sovrn_metrics_vm_port }} -vmalert.proxyURL=http://127.0.0.1:{{ sovrn_metrics_vmalert_port }}
# victorialogs.service.j2 — ExecStart line
ExecStart=/usr/local/sbin/victoria-logs -storageDataPath={{ sovrn_metrics_vl_datadir }} -retentionPeriod={{ sovrn_metrics_vl_retention }} -retention.maxDiskSpaceUsageBytes={{ sovrn_metrics_vl_max_disk }} -httpListenAddr=127.0.0.1:{{ sovrn_metrics_vl_port }} -journald.streamFields=cell,service,_SYSTEMD_UNIT,_HOSTNAME -journald.useRemoteIP=true -vmalert.proxyURL=http://127.0.0.1:{{ sovrn_metrics_vmalert_port }}
# vmalert.service.j2 — ExecStart line (serves BOTH rule types via two -rule files; logs group uses type: vlogs)
ExecStart=/usr/local/sbin/vmalert -rule=/etc/vmalert/rules-cell.yml -rule=/etc/vmalert/rules-logs.yml -datasource.url=http://127.0.0.1:8428 -remoteWrite.url=http://127.0.0.1:8428 -remoteRead.url=http://127.0.0.1:8428 -notifier.url=http://127.0.0.1:9093 -httpListenAddr=127.0.0.1:8880
# alertmanager.yml.j2 — Resend SMTP, recipient from host_vars
global:
  smtp_smarthost: 'smtp.resend.com:587'
  smtp_from: '{{ sovrn_alert_from }}'
  smtp_auth_username: 'resend'
  smtp_auth_password_file: '/etc/sovrn/secrets/resend-api-key'
route: {receiver: 'cell-ops'}
receivers:
  - name: 'cell-ops'
    email_configs:
      - to: '{{ sovrn_alert_recipient }}'
        send_resolved: true
# /etc/sovrn/secrets/resend-api-key (0600, sovrn-owned) is materialized from
# vault-shared.yml sovrn_pds_resend_api_key by the sovrn_secrets role — same
# fleet key already used for ZDS mail; sender domain notify.sovrn.at already verified.
# rules-cell.yml.j2 — metrics alerts (type: prometheus default)
groups:
  - name: cells
    interval: 30s
    rules:
      - alert: CellTargetDown
        expr: 'up{job="vmagent-cell"} == 0'
        for: 2m
        labels: {severity: page}
        annotations: {summary: 'metrics push down for cell {{$labels.cell}}'}
      - alert: CellDiskFullSoon
        expr: 'predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 7*24*3600) < 0'
        for: 30m
        labels: {severity: page}
        annotations: {summary: 'disk on {{$labels.cell}} fills within 7d at current rate'}
      - alert: CellDiskHigh
        expr: '(1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) > 0.85'
        for: 15m
        labels: {severity: warn}
        annotations: {summary: 'disk above 85% on {{$labels.cell}}'}
      - alert: BackupStale
        expr: 'time() - sovrn_backup_marker_mtime_seconds > 2700'
        for: 10m
        labels: {severity: page}
        annotations: {summary: 'backup marker stale on {{$labels.cell}} ({{$labels.marker}})'}
      - alert: TlsExpiring
        expr: 'sovrn_tls_expiry_days < 14'
        for: 1h
        labels: {severity: warn}
        annotations: {summary: 'TLS cert expiring on {{$labels.cell}} in {{$value}}d'}
# rules-logs.yml.j2 — log alerts (NOTE type: vlogs, datasource for this group is VL :9428)
groups:
  - name: cell-logs
    type: vlogs
    interval: 1m
    rules:
      - alert: CellErrorBurst
        expr: 'cell="*" level="error" | stats count() as n by (cell) | filter n > 20'
        for: 5m
        labels: {severity: warn}
        annotations: {summary: 'error burst on {{$labels.cell}}'}
# Caddyfile-infra.j2 — single host infra.sovrn.at; EVERYTHING behind basicauth with the shared fleet secret
infra.sovrn.at {
	basicauth {
		{$METRICS_USER} {$METRICS_HASH}
	}
	# MetricsQL UI + API
	handle /vmui* { reverse_proxy 127.0.0.1:8428 }
	handle /api/* { reverse_proxy 127.0.0.1:8428 }
	handle /metrics { reverse_proxy 127.0.0.1:8428 }
	# VictoriaLogs UI + API (vmui at /select/vmui/)
	handle /select/* { reverse_proxy 127.0.0.1:9428 }
	handle /insert/* { reverse_proxy 127.0.0.1:9428 }
	# vmalert rules view
	handle /vmalert* { reverse_proxy 127.0.0.1:8880 }
	handle / { reverse_proxy 127.0.0.1:8428 }
}
# {$METRICS_USER}=cell, {$METRICS_HASH}=caddy hash-password of sovrn_metrics_password (rendered via vault, never logged).
# AUTH POSTURE (locked): this Caddy basicauth is the ONLY auth layer. Daemons bind
# 127.0.0.1 with NO -httpAuth.* flags by design — do not add them.
  • [ ] Step 2: Gate site.yml so cell roles skip the infra box and vice versa
# deployment/playbooks/site.yml — add at end, and guard cell-only roles
    - name: Central metrics + logs (infra box only)
      ansible.builtin.include_role:
        name: metrics_box
      tags: [metrics-box]
      when: inventory_hostname in groups['metrics']
# + add to each cell-data-plane role task include:
#     when: inventory_hostname not in groups['metrics']
#   (stalwart, zds, caddy-cell vhost, sovrnd, litestream, rclone — infra box gets common + metrics_box only)
  • [ ] Step 3: Syntax-check and converge the infra box
cd deployment && ansible-playbook playbooks/site.yml -l infra.sovrn.at -e mode=update --tags metrics-box --check --diff
just update infra.sovrn.at --tags metrics-box

Expected: check-mode clean; converge green; systemctl is-active victoriametrics victorialogs vmalert alertmanager caddy all active.

  • [ ] Step 4: Verify endpoints + auth
curl -fsS https://infra.sovrn.at/metrics -u "cell:$(secrets decrypt metrics_ingest_password)" | head -5
curl -s -o /dev/null -w '%{http_code}\n' https://infra.sovrn.at/api/v1/query # expect 401 without creds
curl -fsS 'https://infra.sovrn.at/api/v1/query?query=up' -u "cell:$(secrets decrypt metrics_ingest_password)"

Expected: vm_ series present; unauthenticated write/query → 401; vmui loads at https://infra.sovrn.at/vmui/.

  • [ ] Step 5: Commit
jj commit -m "feat(metrics): metrics_box role on infra.sovrn.at (VM+VL+vmalert+AM+Caddy)" deployment/roles/metrics_box deployment/playbooks/site.yml

Task 3: telemetry_edge role — vmagent + Vector + node_exporter on every cell

Files: - Create: deployment/roles/telemetry_edge/tasks/main.yml - Create: deployment/roles/telemetry_edge/handlers/main.yml - Create: deployment/roles/telemetry_edge/templates/scrape.yml.j2 - Create: deployment/roles/telemetry_edge/templates/vmagent.service.j2 - Create: deployment/roles/telemetry_edge/templates/vector.yaml.j2 - Create: deployment/roles/telemetry_edge/templates/vector.service.j2 - Create: deployment/roles/telemetry_edge/templates/node_exporter.service.j2 - Modify: deployment/playbooks/site.yml (include edge role on mail_primary, tag telemetry-edge)

  • [ ] Step 1: Write cell scrape config (assumes /metrics contract from child issue; missing targets fail visibly, not silently)
# scrape.yml.j2 — vmagent scrapes loopback only; cell label injected once
global:
  scrape_interval: 15s
  external_labels: {cell: '{{ inventory_hostname }}'}
scrape_configs:
  - job_name: host
    static_configs: [{targets: ['127.0.0.1:9100'], labels: {service: host}}]
  - job_name: stalwart
    metrics_path: /metrics/prometheus
    static_configs: [{targets: ['127.0.0.1:8080'], labels: {service: stalwart}}]
  - job_name: sovrnd
    metrics_path: /metrics
    static_configs: [{targets: ['127.0.0.1:8081'], labels: {service: sovrnd}}]
  - job_name: caddy
    static_configs: [{targets: ['127.0.0.1:2019'], labels: {service: caddy}}]
  - job_name: vmagent-self
    static_configs: [{targets: ['127.0.0.1:8429'], labels: {service: vmagent}}]
# vmagent.service.j2 — ExecStart line
ExecStart=/usr/local/sbin/vmagent -promscrape.config=/etc/vm/scrape.yml -remoteWrite.url=https://infra.sovrn.at/api/v1/write -remoteWrite.basicAuth.username=cell -remoteWrite.basicAuth.passwordFile=/etc/vm/ingest-password -remoteWrite.tmpDataPath=/var/lib/vmagent-buffer -remoteWrite.maxDiskUsagePerURL=2GiB -extra_label=cell={{ inventory_hostname }} -httpListenAddr=127.0.0.1:8429
# vector.yaml.j2 — journald source, cell+service enrichment, disk-buffered VL sink
sources:
  cell_journald: {type: journald, data_dir: /var/lib/vector, current_boot_only: true}
transforms:
  add_cell:
    type: remap
    inputs: [cell_journald]
    source: |
      .cell = "{{ inventory_hostname }}"
      .service = string!(._SYSTEMD_UNIT)      
sinks:
  central_vl:
    type: elasticsearch
    inputs: [add_cell]
    endpoints: ["https://infra.sovrn.at:443/insert/elasticsearch/"]
    auth: {strategy: basic, user: cell, password: "{PASSWORD_FILE:/etc/vm/ingest-password}"}
    api_version: v8
    compression: gzip
    healthcheck: {enabled: false}
    query: {_msg_field: MESSAGE, _time_field: __REALTIME_TIMESTAMP, _stream_fields: cell,service}
    batch: {max_bytes: 1048576, max_events: 1000, timeout_secs: 2}
    buffer: {type: disk, max_size: 536870912, when_full: block}
# NOTE: Vector runs as root (journald read for ALL units incl. user zds@* units);
# /etc/vm/ingest-password is 0600 root from vault-shared sovrn_metrics_password.
  • [ ] Step 2: Wire edge role into site.yml (cells only)
    - name: Telemetry edge (cells push to infra box)
      ansible.builtin.include_role:
        name: telemetry_edge
      tags: [telemetry-edge]
      when: inventory_hostname in groups['mail_primary']
  • [ ] Step 3: Converge one canary cell first
cd deployment && ansible-playbook playbooks/site.yml -l sovrn.at -e mode=update --tags telemetry-edge --check --diff
just update sovrn.at --tags telemetry-edge
curl -fsS 'https://infra.sovrn.at/api/v1/query?query=up{cell="sovrn.at"}' -u "cell:$(secrets decrypt metrics_ingest_password)"

Expected: up{cell="sovrn.at",service=~"host|stalwart|sovrnd"} present; sovrnd target up==0 until the child-issue /metrics endpoint lands (visible gap, tracked — does not block other targets).

  • [ ] Step 4: Verify logs decompose by cell and service
# In VL vmui (https://infra.sovrn.at/select/vmui/), run:
#   cell="sovrn.at" | stats count() by (service)
# Expect rows for stalwart.service, sovrnd.service, zds@*.service, caddy, vector.
curl -fsS -u "cell:$(secrets decrypt metrics_ingest_password)" \
  'https://infra.sovrn.at/select/logsql/query?query=cell%3D%22sovrn.at%22%20%7C%20stats%20count()%20by%20(service)'

Expected: per-service counts; zds@* user-unit lines present (proves root journald read covers user journals).

  • [ ] Step 5: Roll out to remaining cells, then commit
just update mx1.eu.sovrn.at --tags telemetry-edge
jj commit -m "feat(metrics): telemetry_edge role — vmagent push + Vector journald ship on cells" deployment/roles/telemetry_edge deployment/playbooks/site.yml

Task 4: Stalwart journald unification (kill the file-tracer special case)

Files: - Modify: deployment/roles/stalwart/files/bootstrap-stalwart.sh (tracer section) - Modify: deployment/inventory/group_vars/all/sovrn.yml (retire or repurpose sovrn_stalwart_logdir) - Modify: deployment/roles/stalwart/tasks/install.yml (drop /var/log/stalwart creation if present)

  • [ ] Step 1: Change the Stalwart tracer from file to stdout
# Locate the exact stanza first — pattern from current code:
grep -n 'Log.*path:/var/log/stalwart\|prefix:stalwart' deployment/roles/stalwart/files/bootstrap-stalwart.sh
# Verify stdout-tracer syntax against the pinned source BEFORE editing:
# (Stalwart checkout ~/projects/stalwart, version per docs/05 Source-of-truth line)
grep -rn 'stdout' ~/projects/stalwart/crates/*/src/tracer* 2>/dev/null | head -20

Then replace the file tracer with the stdout equivalent (exact keys per checkout; stdout lines inherit SYSLOG_IDENTIFIER=stalwart via the systemd unit, and Vector maps _SYSTEMD_UNIT=stalwart.service → service="stalwart.service").

  • [ ] Step 2: Converge canary cell, verify no file logs and VL continuity
just update sovrn.at --tags stalwart-install,telemetry-edge
ssh sovrn.at 'ls /var/log/stalwart 2>&1; journalctl -u stalwart --no-pager -n 5'
# In VL vmui: cell="sovrn.at" service="stalwart.service" | stats count() — expect fresh rows post-restart

Expected: /var/log/stalwart absent; journalctl -u stalwart shows live lines; VL query returns post-restart rows (no gap beyond restart window, proving the file-tail removal lost nothing).

  • [ ] Step 3: Commit
jj commit -m "feat(metrics): stalwart logs to journald, retire /var/log/stalwart" deployment/roles/stalwart deployment/inventory/group_vars/all/sovrn.yml

Task 5: Dashboards (vmui shareable queries) + staged-failure acceptance

Files: - Create: docs/runbooks/metrics-box.md (vmui URLs, LogsQL/MetricsQL library, alert-response steps) - Modify: docs/deployment.md (infra box section: group, secrets, retention, UI URLs)

  • [ ] Step 1: Record the dashboard query library (vmui shareable URLs, one UI as agreed)
# docs/runbooks/metrics-box.md — core queries (all scoped by cell=~"$cell" pattern)
- Fleet health: `up{job="vmagent-cell"} by (cell)` — every cell must be 1.
- Disk forecast: `predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 7*24*3600) by (cell)` + current `%`: `(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) by (cell)`.
- Queues: Stalwart queue gauge from `job="stalwart"` (exact name per /metrics/prometheus inventory in Task 3 verification; record here).
- Backup freshness: `time() - sovrn_backup_marker_mtime_seconds by (cell, marker)`.
- TLS: `sovrn_tls_expiry_days by (cell)` (< 14 warns).
- Consolidated logs: `cell=~".*" | stats count() by (cell, service)`; drill-down `cell="<c>" service="<s>" <text> | ...`; live tail in /select/vmui/ with `cell="<c>"`.
- Scrape self-health: `up{job=~"host|stalwart|sovrnd|caddy"} by (cell, service)` — pinpoints which leg of a cell is dark.
  • [ ] Step 2: Staged-failure acceptance (the bug’s acceptance criterion — one cell, one failure at a time)
# 1. target-down: stop vmagent on canary cell, expect CellTargetDown page within ~3 min, resolve on restart
ssh sovrn.at 'sudo systemctl stop vmagent'  # then watch vmalert / alert inbox; then start
# 2. queue-depth: hold SMTP delivery (firewall-drop :25 egress on canary) to grow Stalwart queue, expect queue alert
# 3. backup-freshness: age a marker file (touch -d '2 hours ago' /var/lib/sovrn/health/rclone.ok), expect BackupStale
# 4. disk: DO NOT fill prod disk — validate predict_linear math against 6h history in vmui instead, record screenshot/query
# 5. cert-expiry: lower threshold copy of rule evaluated manually in vmui against current sovrn_tls_expiry_days

Expected: each staged failure fires its matching alert to [email protected] and resolves; results logged in docs/runbooks/metrics-box.md with timestamps.

  • [ ] Step 3: Commit + comment results on the bug
jj commit -m "docs(metrics): metrics-box runbook, dashboards, staged-failure results" docs/runbooks/metrics-box.md docs/deployment.md

Task 6: Child issue + Grafana-deferral note

  • [ ] Step 1: Child issue for the sovrn /metrics endpoint is created separately (agent runs bug agent new --parent 6cae3a0 with the metric list below) — Task 3 scrape config already targets it, so this box plan stays unblocked.
  • [ ] Step 2: Record Grafana deferral: when 5+ cells or shared-dashboard need arises, add Grafana :3000 on infra box with VM+VL datasources; no schema changes needed (labels are already Grafana-friendly).

Self-review

  1. Spec coverage: metrics box on separate host ✓ (Task 2, infra.sovrn.at) · VM + VL instances ✓ · all cells scraping ✓ (Task 3 vmagent push) · log forwarding ✓ (Task 3 Vector) · file-vs-journald investigation ✓ (resolved: journald unification, Task 4; rationale: single pipeline, no double-write, Vector journald source + disk buffer beats journal-upload flakiness) · disk/queue/health dashboard ✓ (Task 5) · consolidated logs by cell+service ✓ (stream fields + Task 5 queries) · vmui vs Grafana evaluation ✓ (vmui day-1, Grafana deferred with trigger) · acceptance staged failure ✓ (Task 5 Step 2).
  2. Placeholder scan: versions resolved via explicit lookup step (Task 1 Step 4) rather than invented; Stalwart stdout syntax verified against pinned checkout before editing (Task 4 Step 1); secret material handled as operator step, never inline.
  3. Type consistency: label contract cell=<inventory_hostname>, service=<job|SYSTEMD_UNIT> used identically in vmagent external_labels, Vector remap, VL stream fields, alert exprs, and dashboard queries.
agent 69c3a3e Sep 16

Implementation complete (12 commits, all spec+quality reviewed, final integration review: Ready to converge). Tasks 1-6 done: inventory+versions (VM 1.152.0, VL 1.52.0, Vector 0.58.0, node_exporter 1.12.1, AM 0.34.0, all checksummed), metrics_box role (VM+VL+dual vmalert+AM+Caddy basicauth-only on infra.sovrn.at), telemetry_edge role (vmagent push + Vector journald ship + node_exporter on cells), Stalwart->journald migration (live cells need manual JMAP x:Tracer migration, must NOT re-bootstrap), runbook docs/runbooks/metrics-box.md + staged-failure procedures, Grafana deferred. Child issues: dfb37b1 (/metrics endpoint), ba06600 (caddy_acme_ca undefined var, pre-existing). OPERATOR NEXT: 1) vault sovrn_metrics_password + bcrypt hash (just provision-metrics-secret) and confirm resend key; 2) DNS A for infra.sovrn.at; 3) triage sovrn.at SSH host-key change (do NOT blindly ssh-keygen -R); 4) just update infra.sovrn.at –tags metrics-box, then one canary cell –tags telemetry-edge; 5) run staged failures, log in runbook section 9. Static verification only — no live converge was possible from here.

agent 66cea9e Oct 6

Status (2026-10-06): the Ansible metrics box from 69c3a3e was superseded by the NixOS metrics role (dd3c64a); infra.sovrn.at was torn down and is rebuilt on the shared netcup box from ~/projects/servers (plan Phase 8). Acceptance (a staged failure on a cell fires the matching alert) is still owed and needs a live cell: plan Phase 9. Cell-side telemetry is a6edca3; dfb37b1 (sovrnd /metrics) is needed before BackupStale/TlsExpiring can fire.

agent 62cca3e Oct 7

Metrics box live on NixOS (2026-10-07, fleet Phase 8): infra.sovrn.at on the shared netcup box (infra.mymood.at), generation 15. LE cert, /healthz 200, UIs behind basic auth, test alert delivered via Resend. Fixed on the way (sovrn qtqttunr): Alertmanager now serves under /am, where Caddy proxies it (was a 404, as under Ansible). Remaining here: a staged failure firing the matching alert from a real cell (Phase 9).

agent 63c4a8e Oct 8

Closing (user, 2026-10-08): acceptance met. The metrics box runs on NixOS (infra.sovrn.at); a staged failure on mx99 fired CellErrorBurst and reached [email protected]; vmui covers every cell, the relay and the shared box from one UI. Still open separately: dfb37b1 (sovrnd exports no sovrn_backup_marker_mtime_seconds / sovrn_tls_expiry_days yet, so BackupStale and TlsExpiring can’t fire).