Stalwart NixOS module on nixpkgs stalwart_0_16 (pinned 0.16.23)

closed
#9f24376 opened by agent Sep 29

Our own NixOS module for Stalwart. nixpkgs services.stalwart targets 0.15.5 and is documented as incompatible with stalwart_0_16.

Sketch (nix/modules/stalwart.nix, option namespace sovrn.services.stalwart): - package = pkgs.stalwart_0_16 with assert package.version == "0.16.23", so a nixpkgs bump that moves Stalwart fails evaluation instead of reaching prod. Bumping is a deliberate edit. (OSS build: features sqlite/postgres/mysql/rocks/s3/redis/azure/nats, no enterprise. Cached on cache.nixos.org for x86_64 and aarch64.) - config.json via environment.etc."stalwart/config.json": DataStore only (sqlite at /var/lib/stalwart/stalwart.db). It holds no secrets, so it is safe in the store. - Service unit: - StateDirectory=stalwart, DynamicUser=no (stalwart user; Litestream and the restore need stable ownership) - AmbientCapabilities=CAP_NET_BIND_SERVICE, systemd hardening - STALWART_RECOVERY_ADMIN from a Colmena key file (EnvironmentFile rendered by a key-ordered oneshot, or LoadCredential) - No apply plan here; that is the plan issue.

Spike first: on 0.16.23, confirm what happens with config.json + an empty sqlite DB, with and without STALWART_RECOVERY_MODE=1. Does it start in bootstrap mode or normal mode with defaults? Which listeners exist by default? Is the recovery admin honored on every start?

Done when the module boots Stalwart on a scratch box and stalwart-cli --url http://127.0.0.1:8080 query Listener works with the recovery admin.

1 Comment

agent 9ffc294 Sep 30

Done (pending review)

nix/modules/stalwart.nix (option sovrn.services.stalwart) is imported by base.nix. No role enables it yet (see finding 1). - Package: pkgs.sovrn.stalwart (nixpkgs stalwart_0_16, version-asserted in nix/overlay.nix). stalwart-cli goes in systemPackages. - config.json: /etc/stalwart/config.json, Sqlite DataStore at /var/lib/stalwart/stalwart.db, generated in the store with restartTriggers. - User and state: static stalwart user; StateDirectory 0750, UMask 0027. - Recovery admin: sovrn-admin by default, the same name as dev. The password comes from the stalwart-recovery-password Colmena key via LoadCredential. A wrapper exports STALWART_RECOVERY_ADMIN from $CREDENTIALS_DIRECTORY, so the key never lands in the unit or an EnvironmentFile. Stalwart reads it from the environment only; there’s no _FILE variant. - Unit: follows upstream (SIGINT, KillMode=process, LimitNOFILE 65536) plus CAP_NET_BIND_SERVICE and a full sandbox (ProtectSystem=strict, SystemCallFilter @system-service ~@privileged, RestrictAddressFamilies, etc.). - Options: environment for extra env (recovery mode, used by ba3c79f) and webuiUrl (read-only, see finding 2).

VM test: nix/tests/stalwart.nix (nix build .#checks.x86_64-linux.stalwart -L) passes, checking: - It runs as stalwart and the DB is owned by stalwart. - stalwart-cli query NetworkListener works with the recovery admin; a wrong password gets 401. - The password is not in the unit’s Environment. - :25 binds without root. - Recovery mode drops the mail listeners and keeps the API.

Findings from the spike (0.16.23): 1. With config.json, an empty DB starts in NORMAL mode (never bootstrap). It seeds and binds the default listeners 25/443/465/993/995/4190/8080: an unconfigured MTA. A host’s first start must be recovery mode plus apply, before normal mode (ba3c79f). Until then no role enables the module. 2. Every start fetches the admin web UI from github.com/stalwartlabs/webui/releases/latest. That’s unpinned (a supply-chain risk), and startup blocks up to 60s when it’s unreachable. file:// URLs are supported, so the module exposes webuiUrl = file://<nixpkgs stalwart_0_16.webui 1.0.11>/webui.zip for the apply plan to set on the Application object’s resourceUrl, with auto-update disabled (b724acb). 3. The recovery admin is honoured in both modes on every start.