systemd orchestration: Stalwart recovery -> converge -> normal
closedWire the plan into boot/deploy as systemd units instead of Ansible’s restart → script → restart dance.
Sketch:
- stalwart-recovery.service: STALWART_RECOVERY_MODE=1, Conflicts=stalwart.service. Runs only when the DB is fresh (ConditionPathExists=! a converged marker, or a DB-empty check).
- stalwart-converge.service (oneshot, RemainAfterExit):
- renders runtime secrets into the plan
- runs stalwart-cli apply --file ...
- runs ReloadSettings
- ordering: After= recovery or the normal server, After=caddy.service for Http01, Wants/After= its key units
- stalwart.service: normal mode, After=stalwart-converge.service on first boot.
- Also re-run converge on colmena apply when the plan hash changes (restartTriggers).
Decision: recovery mode stops MTA/queue, so converging in recovery on every boot/deploy means mail downtime each time. Recommendation: recovery only on a fresh DB; otherwise apply upserts against the running server. Settle it here.
- ACME readiness: a separate check (timer or alert) waiting for a valid covering certificate. It should not block activation.
- Restore interaction: when
sovrn-restore.target(recovery issue) restores a DB, converge runs online, not in recovery mode.
Done when a fresh box reaches a working Stalwart unattended from nixos-anywhere, and reboots and re-deploys don’t re-enter recovery mode.
4 Comments
From 9f24376 (verified on 0.16.23): with config.json present, an empty DB starts in normal mode and binds the default listeners (25/443/465/993/995/4190/8080). The first start on a host must therefore be recovery mode (
sovrn.services.stalwart.environment.STALWART_RECOVERY_MODE = "1", or a separate unit) plusapply, before any normal-mode start. That’s also why no role enables the module yet. Recovery mode binds only the management API (default :8080;STALWART_RECOVERY_MODE_PORTmoves it).From b724acb, for the boot wiring: - Converge command:
sovrn.services.stalwart.plan.converge(stalwart-converge [URL]) renders and applies the host’s plan as the recovery admin. It’s idempotent, so run it on every boot or deploy as you prefer. - First start is slow: the first recovery start on an empty DB takes about 40s (the default web UI Application points at GitHub until the plan pins it). - Network needed: the apply needs outbound network, because creating the AcmeProvider registers with Let’s Encrypt, and a first-boot converge before networking is up will fail on it. Order after network-online and retry on failure. - Restart after apply: listeners bind at startup, so switching to normal mode (a restart) after the apply is required for listener changes. - Online converge: needsReloadSettingsfor settings singletons (e.g. Metrics).Done (pending review)
Decision (user): always start in recovery mode, run
stalwart-converge, then restart in normal mode, on every boot, deploy and restart. The downtime is accepted (senders retry, and the backup relay queues).nix/modules/stalwart.nix, a singlestalwart.service: -ExecStartPre(stalwart-recover-and-converge): - Stalwart in recovery mode onrecoveryPort(18080, off the normal :8080), waiting for/healthz/live(fails fast if it exits) -stalwart-converge, then SIGINT to stop it -ExecStart: normal mode. -ExecStartPost: waits forreadyUrl(127.0.0.1:8080/healthz/ready), so dependents like sovrnd only start against a converged, ready server. - Secrets: the recovery password and every plan secret are passed as systemd credentials (recovery-password,plan-<name>). The stalwart user never reads the key files, e.g.sovrnd-stalwart-passwordis owned by sovrn.stalwart-convergereads the credentials inside the unit and falls back to the key files when run by hand as root. - Ordering: after network-online, since creating or updating the ACME provider talks to the CA, and after/wants the Colmena key units of every file it reads. - Restart triggers: config.json and the converge script. The converge script embeds the plan files and vars, so a plan change on deploy restarts Stalwart, which converges. - Retries:Restart=on-failure,RestartSec=30,StartLimitIntervalSec=0, so a failed converge (e.g. the CA unreachable at boot) keeps retrying instead of leaving the host without mail.TimeoutStartSec=15min. - stalwart-cli environment:CacheDirectory=stalwartfor its schema cache (HOME/XDG_CACHE_HOME), andSSL_CERT_FILEdefaults to the system bundle.Cell role: Stalwart is enabled, with firewall ports 25/465/587/143/993 open (4190 and :8080 stay closed).
VM test
checks.stalwartpasses, all automatic: - The unit converges on first boot in recovery mode (16 created, 0 failed), and :18080 is gone afterwards. - It runs as stalwart; the recovery admin works; neither password appears in the unit’s Environment. - Exactly the cell listeners. - sovrnd is allowed and denied the right things. - The ACME certificate is issued via pebble. -systemctl restart stalwartconverges again with 0 created, and the certificate persists.Timing (VM, no network): first boot about 104s, restart about 82s. Most of it is each instance waiting on lookups that time out without network. Locally with network, a recovery start took about 9s and a normal start about 2s. Not measured on a real cell yet.
Known gaps (other issues): - ACME HTTP-01 on a real cell needs Caddy’s :80 challenge proxy (63eb0d9); until then certificates can’t issue there. - On a first Colmena deploy,
sovrnd-stalwart-passworduploads post-activation (owner sovrn), so the first start fails on the missing credential and succeeds on the 30s retry once the key lands. - The relay role doesn’t enable Stalwart yet (b45a76f).Timing correction (from 63eb0d9): the ~40s per Stalwart instance in the VM tests was DNS timeouts against the base config’s unreachable resolvers, not Stalwart itself. With lookups failing fast in the VM, a whole start (recovery → converge → normal → ready) takes about 21 seconds on a fresh database. Both VM tests now configure no resolvers.