Backup relay (mxb) on NixOS

closed
#b45a76f opened by agent Sep 29

mxb..sovrn.at: second real target. It has no mailboxes, only a transient 5–7 day queue.

Sketch: - Stalwart module + relay.plan.json: :25-only listeners, ACME for the relay hostname, queue TTL. - backup-route-sync service (from #3). It polls every cell’s GET /backup/domains and owns routes, schedule, strategy and the RCPT relaying exception. Config comes from Nix: sovrn_backup_cells becomes a module option; the bearer token is a key. - Caddy with the challenge-proxy-only posture. - Firewall: 22/25/80 only. - Telemetry edge (see the telemetry issue).

Done when the relay accepts mail for a cell’s domain while the cell is down, and delivers it when the cell returns.

6 Comments

agent b04257a Sep 30

From b724acb: the relay needs its own plan, nix/stalwart/relay.plan.json, with a relay composition in nix/stalwart/default.nix: common + relay. It needs no roles.plan.json/sovrnd.plan.json, since the relay has no sovrnd and no tenant accounts. Contents: - AcmeProvider + a Domain for the relay hostname - SystemSettings - listeners (reconcile) :25 only, plus http on 127.0.0.1:8080

backup-route-sync still owns routes, schedule, strategy and relaying at runtime. Decide whether it keeps using the recovery admin or gets its own restricted account (MtaRoute/MtaDeliverySchedule/MtaOutboundStrategy/MtaStageRcpt), the same pattern as sovrnd’s role. Add the composition to checks.stalwart-plan.

agent bc4c5da Oct 6

Update (2026-10-06): the relay runs on the shared netcup box (infra.mymood.at, us-east), deployed from ~/projects/servers, not on a dedicated mxb box. Consequences for the role: - sovrn.relay.hostname option (mxb.eu.sovrn.at); never the machine’s own name. - Caddy is shared: the relay adds an http://<relay hostname> vhost that proxies only the ACME challenge to Stalwart; it doesn’t own 80⁄443. - netcup’s default firewall policy blocked outbound 25/465/587; the rule was deleted and outbound 25 verified. Add an outbound-25 probe to the relay’s health so a re-applied block becomes an alert. - PTR for the box (IPv4 + IPv6) → mxb.eu.sovrn.at before go-live; inbound 25 still untested. - Done-condition gets a VM test (relay + fake cell: queue while down, drain on return); the live check happens when mx99 is rebuilt (plan Phase 9).

agent bf495da Oct 6

Implemented (2026-10-06), pending review

Port decision (user asked whether relay → cell can use an open port like 2525): yes, and it’s better than 25. Stalwart 0.16 treats every port but 25 as submission, and on 25 checks SPF/DMARC/reverse IP against the connecting IP, which for relayed mail is the relay’s. In the default relaxed mode that doesn’t reject (only strict does), but legitimate mail queued during an outage would be recorded as SPF/DMARC failures, fed to the spam filter, and trigger DMARC failure reports. Cells now take relayed mail on a dedicated relay listener (sovrn.cell.relayPort, 2525) that: - requires no auth and offers no SASL (MtaStageAuth override; 587⁄465 unchanged) - keeps SPF/DMARC/reverse-IP off (Stalwart’s non-25 default; the relay checked the real sender on its own :25) - adds a Received header (MtaStageData override) and still runs the spam filter - is firewalled to sovrn.cell.relaySources (the fleet fills it with every sovrn-relay host’s addresses)

The relay therefore needs no outbound port 25 at all.

Commits: wvrkyoyu (cell listener/overrides/firewall, route-sync CELL_PORT), krpvlntz (fixes found by the new test, below), nkprwoko (relay role, nix/stalwart/relay.plan.json, nixosModules.relay, gen-secret).

route-sync account: its own restricted account route-sync@<relay hostname> with a role in the plan (MtaRoute/MtaDeliverySchedule/MtaOutboundStrategy/MtaStageRcpt and the settings reload only), not the recovery admin.

Bugs found by checks.relay (first live run of route-sync and of a non-sovrnd account): - internal/stalwart sent accountId "a" on registry calls, which Stalwart only allows for callers holding impersonate (sovrnd has it, so it never showed). Now omitted; Stalwart resolves it to the caller. sovrnd’s behaviour is unchanged (account filtering only applies without impersonate). - routesync’s MtaDeliverySchedule.retry shape was rejected; the Custom variant is flattened and a List is an index-keyed object.

Tests: new checks.relay (fake cell down, then up: other domains refused, mail queued, delivered on 2525 at the first retry); checks.cell covers the 2525 listener; stalwart-plan includes the relay composition; all checks and Go tests pass.

Known gap (unchanged from Ansible): bounces for mail expiring on the relay follow the strategy fallback (first cell), which won’t relay foreign domains, so they’re lost. Fix later by sending non-cell destinations through the SMTP2GO smarthost.

Deploying to the netcup box and live acceptance: plan Phases 8–9.

agent b9475da Oct 7

Relay live on NixOS (2026-10-07, fleet Phase 8): mxb.eu.sovrn.at on infra.mymood.at. Stalwart LE cert via Caddy’s challenge proxy, STARTTLS verifies (v4+v6), RCPT for any domain 550 5.1.2 (no cells), PTR v4+v6 = mxb.eu.sovrn.at. Blocked: inbound 25 doesn’t reach the box (netcup firewall policy upstream; no SYN arrives). Remaining: that, then Phase 9 (route-sync against mx99, queue-and-drain live).

agent bf475ba Oct 8

Live queue-and-drain passed (fleet Phase 9, 2026-10-08): mx99 (NixOS staging cell, arm64) Stalwart stopped 03:51:02; Gmail -> [email protected] fell back to the pri-20 MX mxb.eu.sovrn.at (relay on infra.mymood.at) at 03:51:55, SPF/DKIM/DMARC pass; queued 03:52:53, first attempt to mx99:2525 refused (v4+v6), rescheduled +5 min, expiry 7 d; Stalwart back 03:53:42; relay retry 03:57:55 on 2525 with STARTTLS, 250 at 03:58:22, queue empty; mx99 ingested it into the mailbox (relay listener 2525) 03:58:51. route-sync had applied the domain from mx99’s /backup/domains (domains:1). backupMx = mxb.eu.sovrn.at is set in sovrn fleet.nix and deployed to mx99.

agent b34a5da Oct 8

Closing (user, 2026-10-08): done. The relay runs on NixOS as mxb.eu.sovrn.at (fleet, infra.mymood.at); live queue-and-drain passed twice on mx99 (Phase 9 test, and the Phase 10 rebuild where route-sync kept the last-known domains and the relay delivered at its next retry).