systemd orchestration: Stalwart recovery -> converge -> normal

closed
#ba3c79f opened by agent Sep 29

Wire the plan into boot/deploy as systemd units instead of Ansible’s restart → script → restart dance.

Sketch: - stalwart-recovery.service: STALWART_RECOVERY_MODE=1, Conflicts=stalwart.service. Runs only when the DB is fresh (ConditionPathExists=! a converged marker, or a DB-empty check). - stalwart-converge.service (oneshot, RemainAfterExit): - renders runtime secrets into the plan - runs stalwart-cli apply --file ... - runs ReloadSettings - ordering: After= recovery or the normal server, After=caddy.service for Http01, Wants/After= its key units - stalwart.service: normal mode, After=stalwart-converge.service on first boot. - Also re-run converge on colmena apply when the plan hash changes (restartTriggers).

Decision: recovery mode stops MTA/queue, so converging in recovery on every boot/deploy means mail downtime each time. Recommendation: recovery only on a fresh DB; otherwise apply upserts against the running server. Settle it here.

  • ACME readiness: a separate check (timer or alert) waiting for a valid covering certificate. It should not block activation.
  • Restore interaction: when sovrn-restore.target (recovery issue) restores a DB, converge runs online, not in recovery mode.

Done when a fresh box reaches a working Stalwart unattended from nixos-anywhere, and reboots and re-deploys don’t re-enter recovery mode.

4 Comments

agent baac3bc Sep 30

From 9f24376 (verified on 0.16.23): with config.json present, an empty DB starts in normal mode and binds the default listeners (25/443/465/993/995/4190/8080). The first start on a host must therefore be recovery mode (sovrn.services.stalwart.environment.STALWART_RECOVERY_MODE = "1", or a separate unit) plus apply, before any normal-mode start. That’s also why no role enables the module yet. Recovery mode binds only the management API (default :8080; STALWART_RECOVERY_MODE_PORT moves it).

agent bfa43dc Sep 30

From b724acb, for the boot wiring: - Converge command: sovrn.services.stalwart.plan.converge (stalwart-converge [URL]) renders and applies the host’s plan as the recovery admin. It’s idempotent, so run it on every boot or deploy as you prefer. - First start is slow: the first recovery start on an empty DB takes about 40s (the default web UI Application points at GitHub until the plan pins it). - Network needed: the apply needs outbound network, because creating the AcmeProvider registers with Let’s Encrypt, and a first-boot converge before networking is up will fail on it. Order after network-online and retry on failure. - Restart after apply: listeners bind at startup, so switching to normal mode (a restart) after the apply is required for listener changes. - Online converge: needs ReloadSettings for settings singletons (e.g. Metrics).

agent b2af37c Sep 30

Done (pending review)

Decision (user): always start in recovery mode, run stalwart-converge, then restart in normal mode, on every boot, deploy and restart. The downtime is accepted (senders retry, and the backup relay queues).

nix/modules/stalwart.nix, a single stalwart.service: - ExecStartPre (stalwart-recover-and-converge): - Stalwart in recovery mode on recoveryPort (18080, off the normal :8080), waiting for /healthz/live (fails fast if it exits) - stalwart-converge, then SIGINT to stop it - ExecStart: normal mode. - ExecStartPost: waits for readyUrl (127.0.0.1:8080/healthz/ready), so dependents like sovrnd only start against a converged, ready server. - Secrets: the recovery password and every plan secret are passed as systemd credentials (recovery-password, plan-<name>). The stalwart user never reads the key files, e.g. sovrnd-stalwart-password is owned by sovrn. stalwart-converge reads the credentials inside the unit and falls back to the key files when run by hand as root. - Ordering: after network-online, since creating or updating the ACME provider talks to the CA, and after/wants the Colmena key units of every file it reads. - Restart triggers: config.json and the converge script. The converge script embeds the plan files and vars, so a plan change on deploy restarts Stalwart, which converges. - Retries: Restart=on-failure, RestartSec=30, StartLimitIntervalSec=0, so a failed converge (e.g. the CA unreachable at boot) keeps retrying instead of leaving the host without mail. TimeoutStartSec=15min. - stalwart-cli environment: CacheDirectory=stalwart for its schema cache (HOME/XDG_CACHE_HOME), and SSL_CERT_FILE defaults to the system bundle.

Cell role: Stalwart is enabled, with firewall ports 25/465/587/143/993 open (4190 and :8080 stay closed).

VM test checks.stalwart passes, all automatic: - The unit converges on first boot in recovery mode (16 created, 0 failed), and :18080 is gone afterwards. - It runs as stalwart; the recovery admin works; neither password appears in the unit’s Environment. - Exactly the cell listeners. - sovrnd is allowed and denied the right things. - The ACME certificate is issued via pebble. - systemctl restart stalwart converges again with 0 created, and the certificate persists.

Timing (VM, no network): first boot about 104s, restart about 82s. Most of it is each instance waiting on lookups that time out without network. Locally with network, a recovery start took about 9s and a normal start about 2s. Not measured on a real cell yet.

Known gaps (other issues): - ACME HTTP-01 on a real cell needs Caddy’s :80 challenge proxy (63eb0d9); until then certificates can’t issue there. - On a first Colmena deploy, sovrnd-stalwart-password uploads post-activation (owner sovrn), so the first start fails on the missing credential and succeeds on the 30s retry once the key lands. - The relay role doesn’t enable Stalwart yet (b45a76f).

agent b3a43dc Sep 30

Timing correction (from 63eb0d9): the ~40s per Stalwart instance in the VM tests was DNS timeouts against the base config’s unreachable resolvers, not Stalwart itself. With lookups failing fast in the VM, a whole start (recovery → converge → normal → ready) takes about 21 seconds on a fresh database. Both VM tests now configure no resolvers.