Evaluate NixOS and Colmena

closed
#ebce532 opened by BT Sep 29

The current deployment uses Ansible for provisioning three kinds of hosts: a cell (Stalwart, sovrnd, zds, backups, metrics, log forwarding, etc.), a backup relay (stalwart to queue mail when the primary cell is down), and a metrics box (Victoria logs and metrics).

I want to evaluate switching from Ansible’s imperative provisioning to NixOS with Colmena. The Ansible configuration is difficult to understand, requires extensive defensive programming in the Justfile to use correctly, and often very brittle. It takes over an hour to run in CI due to building Stalwart from source, and uses custom Python scripts to bootstrap the Stalwart server.

The intent is to create an overall tracking issue in bug with child issues that outline a step-by-step approach to improving the build and deployment process.

  1. Use Stalwart from Nixpkgs rather than building from source: rely on the nixpkgs version of Stalwart with sqlite enabled and pin to a known version number to prevent minor version bumps < 1.0.0 from breaking prod. Building from source takes over an hour and a Nix binary cache should be much faster and support multiple architectures (amd64 + arm64)

  2. Follow the guidance in the Stalwart docs for declarative deployments: Provision an /etc/stalwart/config.json that sets the store to Sqlite and start the service in recovery mode using a pre-generated secret. Use stalwart-cli to bootstrap the service rather than custom python. Use stalwart-cli apply and snapshot to converge on a known-good config on every boot.

  3. Use systemd dependencies, one-shot services, and service ordering: Instead of imperative Ansible steps, use a set of systemd services that order themselves via deps and service ordering to start Stalwart in recovery, apply a known config, then restart the service in prod. In a backup recovery scenario, activate services to restore databases and data, then bring up production services.

  4. Integrate the secrets (~/projects/secrets) tool with Colmena’s secrets functionality: Use Colmena’s deployment.keys..keyCommand to run secrets decrypt <name> to simplify getting encrypted secrets on the boxes. This is easier to reason about than Ansible’s opaque ansible-vault files where everything is encrypted.

Questions:

  1. Can colmena provision the three types of boxes, including the imperative steps required to start Stalwart in recovery, configure it, and create appropriate service accounts?

  2. What are the bite-sized issues that need to be tackled to migrate from Ansible to a NixOS deployment? The first must be provisioning NixOS on hetzner using nixos-anywhere, since they do not provide a NixOS image. Each should be sketched out in a separate issue so that further architecture work can be decided during the migration to Nix.

2 Comments

agent e3bec8e Sep 29

Evaluation plan: Ansible → NixOS + Colmena

Findings (verified 2026-09-29)

Stalwart in nixpkgs - pkgs.stalwart is pinned to 0.15.5, and services.stalwart targets that version. nixpkgs also ships stalwart_0_16 = 0.16.23, the same version as nix/pkgs/stalwart.nix. nixpkgs warns: “stalwart_0_16 is not compatible with services.stalwart at this time.” So we use the nixpkgs package but write our own small NixOS module. - stalwart_0_16 builds with buildNoDefaultFeatures = true and features sqlite postgres mysql rocks s3 redis azure nats. There is no enterprise, so it is still a true OSS/AGPL build, just with extra store backends. - Both x86_64-linux and aarch64-linux outputs are on cache.nixos.org, so there is no hour-long Rust build. - Pinning: the nixpkgs flake.lock plus an assert pkg.version == "0.16.23" in the module. A nixpkgs bump that moves Stalwart then fails evaluation and does not reach prod silently. - stalwart-cli 1.0.12 is in nixpkgs. It has apply and snapshot, reads auth from the environment (STALWART_URL/USER/PASSWORD/TOKEN), and supports apply --dry-run.

Stalwart declarative flow (https://stalw.art/docs/configuration/declarative-deployments/) - config.json holds only the DataStore (sqlite path) and contains no secrets, so it is safe in the Nix store. - Flow: STALWART_RECOVERY_MODE=1 + STALWART_RECOVERY_ADMIN → stalwart-cli apply (NDJSON plan: upsert/reconcile/update/create/destroy, in dependency-aware order) → restart without the recovery env. - Recovery mode stops the MTA, queue and tasks, and serves only the management API on :8080. - reconcile + scope replaces the hand-written listener WANT/DROP logic in bootstrap-stalwart.sh.

Q1: Can Colmena provision all three box types, including the imperative Stalwart steps?

Yes. Colmena itself stays declarative. Every imperative step becomes an ordered systemd oneshot in the NixOS config, and Colmena only pushes closures and keys. Caveats that the child issues must handle:

  1. Keys vanish on reboot by default. deployment.keys upload to /run/keys (tmpfs). An unattended reboot would leave Stalwart, sovrnd and litestream without secrets until the next colmena upload-keys. For mail boxes, set destDir to a persistent root-only path (e.g. /var/lib/sovrn/keys) or accept a deploy after every reboot. Order services on Colmena’s generated <name>-key.service units.
  2. Secrets must never enter the Nix store. Anything that interpolates a secret into the apply plan (service-account passwords, relay creds) is rendered at runtime by the oneshot (e.g. jq --rawfile from the key file). It is never built as a derivation.
  3. Pre-generate the recovery admin with secrets instead of generating it on the box and fetching it back. This removes ensure-bootstrap-secrets, persist-bootstrap-secrets, _vault_lib.py and the sentinel/slurp retry logic entirely.
  4. Recovery mode on every boot means MTA downtime on every boot and deploy. Recommendation: use recovery mode only when the DB is fresh (first boot, or after a restore that found no DB). Otherwise run apply (upserts) against the running server and follow with ReloadSettings. To be decided in the orchestration issue.
  5. ACME stays an ordering dependency. It needs Caddy’s challenge proxy, so the converge unit runs After=caddy.service. Waiting for the cert becomes a separate check or alert rather than blocking activation.
  6. Registry quirks found in bootstrap-stalwart.sh must be re-verified in plan form: Map fields are sets on the wire; the Metrics singleton needs ReloadSettings; AcmeProvider has no name; bootstrap-mode Bootstrap/get returns stale data. Start from stalwart-cli snapshot of mx99, not from scratch.
  7. sovrnd mutates runtime state. It runs systemctl --user for zds@<slug> and writes the Caddy maps. This works on NixOS (users.users.sovrn.linger = true, systemd.user.services."zds@"), but the Caddyfile must keep importing mutable map files outside the store.
  8. Hetzner has no NixOS image. Use nixos-anywhere + disko (kexec needs about 1 GB+ RAM). IPv6 on Hetzner Cloud needs static config per host. ufw becomes networking.firewall, and the resolved drop-in becomes services.resolved.
  9. Whole classes of defensive code disappear: patchelf/PT_INTERP fixes, glibc-ceiling gates, controller-built binary prechecks, and hostname-regex role dispatch in the Justfile (a Colmena node’s role is its imported modules).

Q2: Proposed child issues (in priority order)

  1. Provision NixOS on Hetzner with nixos-anywhere + disko. Add flake.nix outputs for a Colmena hive and a base module: sovrn user, SSH, firewall, resolved (Hetzner + Quad9), Hetzner networking including IPv6, amd64 + arm64. Done when a throwaway box goes from nixos-anywhere --flake .#scratch to a clean colmena apply --on scratch.
  2. Colmena keys ↔ secrets integration. keyCommand = [ "secrets" "decrypt" "sovrn/<host>/<name>" ], persistent destDir, owner/group/permissions per key, <name>-key.service ordering. One-off migration of the ansible-vault contents (shared, per-host, per-cell) into the secrets store under a naming scheme. Decide how CI or a deployer gets the age identity.
  3. Package sovrnd (buildGoModule) and backup-route-sync. Add them as flake outputs next to zds. Enforce gen-check inputs in the derivation. Decide on a binary cache for our own packages (attic/cachix vs building on the deployer).
  4. Pilot: metrics box (infra.sovrn.at) on NixOS. VictoriaMetrics, VictoriaLogs, vmalert, Alertmanager, Caddy infra vhost, metrics-healthz, using nixpkgs modules and version-checked against the current pins. It holds the least state, so it is the lowest-risk real cutover.
  5. Stalwart NixOS module (sovrn.services.stalwart). stalwart_0_16 with a version assert, config.json via environment.etc, user/dirs/hardening, CAP_NET_BIND_SERVICE, recovery admin from a key file. Spike: confirm 0.16.23 behavior with config.json + empty sqlite + recovery mode.
  6. Stalwart declarative plan. snapshot mx99, then curate it into Nix-generated NDJSON: tracer, listeners (reconcile, cell set vs relay :25-only set), metrics, AcmeProvider, domain cert management, SystemSettings. Inject secrets at runtime. apply --dry-run runs as a flake check. This replaces bootstrap-stalwart.sh.
  7. systemd orchestration for Stalwart. stalwart-recovery.service (Conflicts=stalwart), stalwart-converge.service oneshot, then stalwart.service; fresh-DB detection vs online apply; ReloadSettings; ACME readiness check. Also a dedicated sovrnd service account (API key or scoped admin) instead of sharing the recovery admin.
  8. Backup relay (mxb) on NixOS. Stalwart module + relay plan + backup-route-sync service. It holds no mailboxes, so it is the second real target.
  9. Cell data plane. ZDS (linger, zds@ user template, dirs), sovrnd service + sovrn.toml from Nix (secrets by file path), Caddy with on-demand TLS/ask + runtime maps, did.json.
  10. Backups. Litestream + freshness timer, rclone blob sync + cell-config-sync, health markers under /var/lib/sovrn/health.
  11. Telemetry edge. vmagent, Vector (journald → VictoriaLogs), node_exporter on cells and the relay.
  12. Recovery as systemd. A sovrn-restore.target of oneshots ordered Before= the services: litestream restore-if-absent, integrity checks, re-own, blob sync-down, Caddy storage, Stalwart DB. Gate it with a flag or specialisation. A freshen-backups oneshot replaces recover-backup.yml. Rehearse on a scratch box.
  13. Cutover and decommission. Migrate mx99 via the recovery path (a restore is the migration). Reduce the Justfile deploy recipes to thin colmena wrappers. Delete deployment/ Ansible, the Python vault scripts, just build-stalwart/nix/pkgs/stalwart.nix. Rewrite docs/deployment.md and the runbooks.

Open decisions

  • Recovery-mode converge only on a fresh DB, or on every boot (caveat 4)?
  • Persistent key destDir, or accept a re-deploy after every reboot (caveat 1)?
  • Nixpkgs channel for hosts: nixos-25.xx stable + overlay for Stalwart, or the current devenv-nixpkgs/rolling?
  • Who deploys: laptop only, or CI with its own age identity?
  • arm64 (Hetzner CAX) for any box type?
agent e0b5cee Oct 8

Closing (user, 2026-10-08): evaluated and adopted: the migration epic 3beadb2 moved every host to NixOS + Colmena (deployed from the fleet repository).