Evaluate NixOS and Colmena
closedThe current deployment uses Ansible for provisioning three kinds of hosts: a cell (Stalwart, sovrnd, zds, backups, metrics, log forwarding, etc.), a backup relay (stalwart to queue mail when the primary cell is down), and a metrics box (Victoria logs and metrics).
I want to evaluate switching from Ansible’s imperative provisioning to NixOS with Colmena. The Ansible configuration is difficult to understand, requires extensive defensive programming in the Justfile to use correctly, and often very brittle. It takes over an hour to run in CI due to building Stalwart from source, and uses custom Python scripts to bootstrap the Stalwart server.
The intent is to create an overall tracking issue in bug with child issues that outline a step-by-step approach to improving the build and deployment process.
Use Stalwart from Nixpkgs rather than building from source: rely on the nixpkgs version of Stalwart with sqlite enabled and pin to a known version number to prevent minor version bumps < 1.0.0 from breaking prod. Building from source takes over an hour and a Nix binary cache should be much faster and support multiple architectures (amd64 + arm64)
Follow the guidance in the Stalwart docs for declarative deployments: Provision an /etc/stalwart/config.json that sets the store to Sqlite and start the service in recovery mode using a pre-generated secret. Use
stalwart-clito bootstrap the service rather than custom python. Use stalwart-cliapplyandsnapshotto converge on a known-good config on every boot.Use systemd dependencies, one-shot services, and service ordering: Instead of imperative Ansible steps, use a set of systemd services that order themselves via deps and service ordering to start Stalwart in recovery, apply a known config, then restart the service in prod. In a backup recovery scenario, activate services to restore databases and data, then bring up production services.
Integrate the
secrets(~/projects/secrets) tool with Colmena’s secrets functionality: Use Colmena’s deployment.keys..keyCommand to run secrets decrypt <name>to simplify getting encrypted secrets on the boxes. This is easier to reason about than Ansible’s opaque ansible-vault files where everything is encrypted.
Questions:
Can colmena provision the three types of boxes, including the imperative steps required to start Stalwart in recovery, configure it, and create appropriate service accounts?
What are the bite-sized issues that need to be tackled to migrate from Ansible to a NixOS deployment? The first must be provisioning NixOS on hetzner using nixos-anywhere, since they do not provide a NixOS image. Each should be sketched out in a separate issue so that further architecture work can be decided during the migration to Nix.
2 Comments
Evaluation plan: Ansible → NixOS + Colmena
Findings (verified 2026-09-29)
Stalwart in nixpkgs -
pkgs.stalwartis pinned to 0.15.5, andservices.stalwarttargets that version. nixpkgs also shipsstalwart_0_16= 0.16.23, the same version asnix/pkgs/stalwart.nix. nixpkgs warns: “stalwart_0_16is not compatible withservices.stalwartat this time.” So we use the nixpkgs package but write our own small NixOS module. -stalwart_0_16builds withbuildNoDefaultFeatures = trueand featuressqlite postgres mysql rocks s3 redis azure nats. There is noenterprise, so it is still a true OSS/AGPL build, just with extra store backends. - Bothx86_64-linuxandaarch64-linuxoutputs are on cache.nixos.org, so there is no hour-long Rust build. - Pinning: the nixpkgs flake.lock plus anassert pkg.version == "0.16.23"in the module. A nixpkgs bump that moves Stalwart then fails evaluation and does not reach prod silently. -stalwart-cli1.0.12 is in nixpkgs. It hasapplyandsnapshot, reads auth from the environment (STALWART_URL/USER/PASSWORD/TOKEN), and supportsapply --dry-run.Stalwart declarative flow (https://stalw.art/docs/configuration/declarative-deployments/) -
config.jsonholds only the DataStore (sqlite path) and contains no secrets, so it is safe in the Nix store. - Flow:STALWART_RECOVERY_MODE=1+STALWART_RECOVERY_ADMIN→stalwart-cli apply(NDJSON plan: upsert/reconcile/update/create/destroy, in dependency-aware order) → restart without the recovery env. - Recovery mode stops the MTA, queue and tasks, and serves only the management API on :8080. -reconcile+scopereplaces the hand-written listener WANT/DROP logic inbootstrap-stalwart.sh.Q1: Can Colmena provision all three box types, including the imperative Stalwart steps?
Yes. Colmena itself stays declarative. Every imperative step becomes an ordered systemd oneshot in the NixOS config, and Colmena only pushes closures and keys. Caveats that the child issues must handle:
deployment.keysupload to/run/keys(tmpfs). An unattended reboot would leave Stalwart, sovrnd and litestream without secrets until the nextcolmena upload-keys. For mail boxes, setdestDirto a persistent root-only path (e.g./var/lib/sovrn/keys) or accept a deploy after every reboot. Order services on Colmena’s generated<name>-key.serviceunits.jq --rawfilefrom the key file). It is never built as a derivation.secretsinstead of generating it on the box and fetching it back. This removesensure-bootstrap-secrets,persist-bootstrap-secrets,_vault_lib.pyand the sentinel/slurp retry logic entirely.apply(upserts) against the running server and follow withReloadSettings. To be decided in the orchestration issue.After=caddy.service. Waiting for the cert becomes a separate check or alert rather than blocking activation.bootstrap-stalwart.shmust be re-verified in plan form: MapReloadSettings; AcmeProvider has noname; bootstrap-modeBootstrap/getreturns stale data. Start fromstalwart-cli snapshotof mx99, not from scratch.systemctl --userforzds@<slug>and writes the Caddy maps. This works on NixOS (users.users.sovrn.linger = true,systemd.user.services."zds@"), but the Caddyfile must keep importing mutable map files outside the store.networking.firewall, and the resolved drop-in becomesservices.resolved.Q2: Proposed child issues (in priority order)
flake.nixoutputs for a Colmena hive and abasemodule: sovrn user, SSH, firewall, resolved (Hetzner + Quad9), Hetzner networking including IPv6, amd64 + arm64. Done when a throwaway box goes fromnixos-anywhere --flake .#scratchto a cleancolmena apply --on scratch.secretsintegration.keyCommand = [ "secrets" "decrypt" "sovrn/<host>/<name>" ], persistentdestDir, owner/group/permissions per key,<name>-key.serviceordering. One-off migration of the ansible-vault contents (shared, per-host, per-cell) into thesecretsstore under a naming scheme. Decide how CI or a deployer gets the age identity.gen-checkinputs in the derivation. Decide on a binary cache for our own packages (attic/cachix vs building on the deployer).sovrn.services.stalwart).stalwart_0_16with a version assert,config.jsonviaenvironment.etc, user/dirs/hardening,CAP_NET_BIND_SERVICE, recovery admin from a key file. Spike: confirm 0.16.23 behavior with config.json + empty sqlite + recovery mode.snapshotmx99, then curate it into Nix-generated NDJSON: tracer, listeners (reconcile, cell set vs relay :25-only set), metrics, AcmeProvider, domain cert management, SystemSettings. Inject secrets at runtime.apply --dry-runruns as a flake check. This replacesbootstrap-stalwart.sh.stalwart-recovery.service(Conflicts=stalwart),stalwart-converge.serviceoneshot, thenstalwart.service; fresh-DB detection vs online apply;ReloadSettings; ACME readiness check. Also a dedicated sovrnd service account (API key or scoped admin) instead of sharing the recovery admin.zds@user template, dirs), sovrnd service +sovrn.tomlfrom Nix (secrets by file path), Caddy with on-demand TLS/ask + runtime maps, did.json./var/lib/sovrn/health.sovrn-restore.targetof oneshots orderedBefore=the services: litestream restore-if-absent, integrity checks, re-own, blob sync-down, Caddy storage, Stalwart DB. Gate it with a flag or specialisation. A freshen-backups oneshot replacesrecover-backup.yml. Rehearse on a scratch box.colmenawrappers. Deletedeployment/Ansible, the Python vault scripts,just build-stalwart/nix/pkgs/stalwart.nix. Rewritedocs/deployment.mdand the runbooks.Open decisions
destDir, or accept a re-deploy after every reboot (caveat 1)?nixos-25.xxstable + overlay for Stalwart, or the currentdevenv-nixpkgs/rolling?Closing (user, 2026-10-08): evaluated and adopted: the migration epic 3beadb2 moved every host to NixOS + Colmena (deployed from the fleet repository).