Epic: migrate provisioning from Ansible to NixOS + Colmena

closed
#3beadb2 opened by agent Sep 29

Tracking issue for migrating sovrn provisioning from Ansible (deployment/) to NixOS + Colmena. The evaluation and findings are in ebce532 (comment e3bec8e).

Child issues are listed in the order to tackle them. Each is a sketch; architecture gets decided inside each issue as the migration goes.

Box types: - cell (mxN..sovrn.at): Stalwart, sovrnd, ZDS, Caddy, Litestream, rclone, telemetry edge - backup relay (mxb..sovrn.at): Stalwart with a :25-only relay, backup-route-sync, telemetry edge - metrics (infra.sovrn.at): VictoriaMetrics, VictoriaLogs, vmalert, Alertmanager, Caddy

Note: mx99.eu.sovrn.at is a smoke-test box only. There is no production data to migrate.

Open decisions (settle them in the child issue that needs them): - Recovery-mode converge only on a fresh DB, or on every boot? (ba3c79f) - Persistent Colmena key destDir, or re-deploy after every reboot? (4ce0593) - Host nixpkgs channel: stable nixos-25.xx + a Stalwart pin, or devenv-nixpkgs/rolling? (bc4b527) - Who deploys: laptop only, or CI with its own age identity? (4ce0593) - arm64 (Hetzner CAX) for any box type? (bc4b527)

Children, in order: 1. bc4b527: NixOS on Hetzner (nixos-anywhere + disko) 2. 4ce0593: Colmena keys via secrets 3. d6ee5d9: Nix packages for sovrnd / backup-route-sync 4. dd3c64a: Pilot: metrics box 5. 9f24376: Stalwart NixOS module 6. b724acb: Stalwart bootstrap as a checked-in JSON apply plan 7. ba3c79f: systemd orchestration for Stalwart 8. b45a76f: Backup relay 9. 63eb0d9: Cell data plane 10. 0e32e5a: Backups 11. a6edca3: Telemetry edge 12. 1badc33: Recovery as systemd 13. c3b0d04: Cutover + decommission

2 Comments

agent 34bde4a Oct 6

Update (2026-10-06): deployments move to a shared Colmena fleet in ~/projects/servers that composes several projects per host. The remaining path, mapped to these children, is in ~/projects/servers/docs/fleet-migration-plan.md (Phases 4–11):

  1. ae42c89 fleet integration + key rotation (new child)
  2. b45a76f relay
  3. a6edca3 telemetry edge
  4. 0e32e5a backups
  5. metrics, relay and edge onto the shared box (infra.mymood.at, netcup)
  6. mx99 rebuilt by the fleet; live acceptance for 63eb0d9, a6edca3, 0e32e5a, b45a76f/d43eebe, 6cae3a0; CI staging from the fleet
  7. 1badc33 recovery (or an explicit decision to delete Ansible without it while there’s no production data)
  8. c3b0d04 decommission

Box layout changed: no dedicated metrics or relay box. Both run on the netcup box (us-east) next to moods; the relay keeps the name mxb.eu.sovrn.at (it backs up the EU cells from a different cloud and region on purpose). Cells stay on their own hosts.

agent 31bceca Oct 8

Epic complete (2026-10-08): sovrn’s provisioning moved from Ansible to NixOS + Colmena, deployed from the fleet (~/projects/servers, docs/fleet-migration-plan.md Phases 4-11). Metrics box and backup relay run on the shared box; mx99 is a NixOS staging cell, accepted live (signup, mail, alerts, backups, relay drain) and rebuilt from its backups in a recovery rehearsal. Remaining children are follow-ups, not migration work: af17761 (nixos-26.11), edbcac6 (forwarding alerts), dc80d76 (Stalwart ASN/Geo download); related: dfb37b1 (sovrnd backup/TLS metrics).