Backup MX (pri-20): store-and-forward relay with postmaster route records
closedGoal
Stand up a single pri-20 backup MX: a Stalwart store-and-forward relay that queues inbound mail while a cell is unreachable and drains when it returns (ADR-0009 D28). No recipient validation in v1.
Decisions locked in discussion (2026-09-18)
- Routing source (option b): new postmaster-authored public record
at.sovrn.domain.routein the per-domain postmaster’s own repo (postmaster.at.<domain>, ADR-0009 D32),key: literal:self. Fields:domain,mxHost,cell,status(active|retired, extensible),updatedAt. Backup jetstreams the collection, filtersauthorDid == postmaster DID, and verifiesresolveHandle(postmaster.at.<domain>) == authorplus domain-suffix match (spoof-proof, no cell coupling). - Rejected: overloading user-authored
at.sovrn.mail.service(untrusted input, per-user cardinality, coverage gap before first login). - Saga: domain setup becomes atomic-ish — Stalwart
EnsureInactive→ DKIM (made idempotent) → DB row (verifying, TID minted upfront) → ZDSSupervisor.Activate→ postmaster account (invite via admin token) → mappingputRecordlast (externally-visible moment). Failure → reverse compensations, best-effort, original error preserved. Postmaster-account rollback is best-effort (deactivate + abandon handle; retry reuses handle via name-taken backstop). - Durability: synchronous compensations plus stateless orphan sweep (no new journal table initially): Stalwart domains without DB rows, stale
verifyingrows with no mapping record,retiringZDS instances past TTL. - Postmaster credentials: one shared pre-provisioned password per cell (new vault var → 0600 file →
[postmaster] passwordfile, same pattern asstalwart-secret). Verified against ZDS source: no admin impersonation exists (repo writes require the account’s own bearer; admin token gates only invites/takedown/sessions), and ZDS has no password-change endpoint, so per-domain random passwords would need unsustainable vault round-trips and still wouldn’t rotate. Envelope-encryption (option 2) deferred to D32 KMS-per-authority work. - Status field: domain-level
statusin the route record covers backup needs (active|retired; retire writesretiredbefore teardown). Account-level enforcement (unpaid → block IMAP/JMAP, spam → block outgoing) stays Stalwart-side; record-status-vs-labeler is a separate spike, not this work. - Known v1 gap (accepted): cross-cell duplicate
domain.createbefore first login is possible (per-cell DBs); login redirect + create-only service record resolve conflicts after first login. A read-only guard (check public route records indomain.create→ 409 + redirect hint) is a fast follow.
Non-goals
Recipient validation on the backup, pri-30, per-domain passwords/KMS, the labeler spike, the cross-cell create guard (fast follow).
15 Comments
Backup MX Implementation Plan
Goal: A single pri-20 Stalwart relay queues mail for unreachable cells and drains on recovery, driven by postmaster-authored
at.sovrn.domain.routerecords with atomic-ish domain setup.Architecture: Cell setup runs a saga (Stalwart → DKIM → DB → ZDS → postmaster account → mapping record last) with reverse compensations plus an orphan sweep; the backup box is a new inventory group running Stalwart with per-cell
Relayroutes and a 5–7d queue, learning domain→cell mappings from the public jetstream.Tech Stack: Go (sovrnd), Stalwart 0.16.x Registry API, ZDS XRPC, Ansible/systemd, jetstream firehose.
Task 1:
at.sovrn.domain.routelexicon + types + validationFiles: - Create:
lexicons/at/sovrn/domain/route.json- Modify:api/sovrn/(regenerate viacmd/sovrn-lexgen— check Justfile/Makefile lexgen target) - Create:internal/postmaster/route.go(record type + validation) - Test:internal/postmaster/route_test.goRecord shape (
key: literal:self, in postmaster’s public repo):Validation rules in
internal/postmaster/route.go(mirrorinternal/pdslifecycle/handles.gostyle — sentinel errors, table tests): -domain,mxHost,celllowercased, non-empty, hostname-label rules via existingcheckLabel-equivalent;domainmust equal the postmaster handle suffix (postmaster.at.<domain>). -statusmust beactive|retired; unknown values rejected (forward-compat: fail closed, matchingcellmatch.goposture). - Author check helper:VerifyAuthor(ctx, domain, authorDID, resolveHandle)— resolvespostmaster.at.<domain>and requires equality withauthorDID.internal/postmaster/route_test.go): table test forValidate(valid record passes; bad hostname, unknown status, domain/handle-suffix mismatch fail) andVerifyAuthorwith a fake resolver (matching DID passes, mismatch fails).go test ./internal/postmaster/ -run TestRoute -v— Expected: FAIL (package undefined).route.goimplementing the rules above.go test ./internal/postmaster/ -v— Expected: PASS.jj commit -m "feat: at.sovrn.domain.route lexicon and validation" lexicons/at/sovrn/domain/route.json api/sovrn internal/postmasterTask 2: Saga prerequisites — idempotent DKIM + upfront TID
Files: - Modify:
internal/dkim/dkim.go(Generategains query-first reuse) - Modify:internal/appview/provision.go(TID minted before Stalwart calls) - Test:internal/dkim/dkim_test.go,internal/appview/provision testsToday
dkim.Generatemints a freshDkimSignatureper call — saga retry would stack signatures. Change: queryDkimSignaturebydomainIdfirst; if present, read back its public half and return without creating (sameKeyderivation as lines 73–96); else current create path.TID: move
id = store.NewID()inProvisionDomain(provision.go:145) above the StalwartEnsurecall and thread it through, so crash-retry reuses the domain name as the idempotency key instead of orphaning under a new TID.x:DkimSignature/setcalls and correctKeyreturn. Provision: test asserting the DB row ID is fixed before/after a Stalwart retry (no duplicateCreateDomainIDs).go test ./internal/dkim/ ./internal/appview/ -v— Expected: FAIL.Generate+ TID hoist.go test ./internal/dkim/ ./internal/appview/ -v— Expected: PASS.jj commit -m "feat: idempotent DKIM generate and upfront domain TID" internal/dkim internal/appviewTask 3: Per-cell postmaster password secret plumbing
Files: - Modify:
deployment/inventory/host_vars/<cell>/vault.ymlpattern (document, operator step — never paste real values) - Create/modify Ansible task: extenddeployment/roles/zds/tasks/secret.ymlpattern with apostmaster-passwordfile task, or a new task indeployment/roles/sovrn_secrets/tasks/main.yml- Modify:deployment/roles/sovrnd/templates/sovrn.toml.j2([postmaster] passwordfile),config.go(Postmasterstruct +bindEnv+ValidateSecrets),deployment/scripts/provision-zds-secrets(generate 32-hex alongside existing keys)New vault var
sovrn_postmaster_password(per-cell, 32 hex like the Stalwart credential) → 0600 file/etc/sovrn/secrets/postmaster-password→sovrn.toml [postmaster] passwordfile.sovrndrefuses to start the domain-setup path when the file is missing (same posture asstalwart.secretfile,config.go:336).config_test.go:ValidateSecretswithoutpostmaster.passwordfilefails mentioning it (mirror the existingrelay.apisecretfiletest,config_test.go:60-82).go test . -run TestValidateSecrets -v— Expected: FAIL.provision-zds-secretskey spec entry.go test . -v— Expected: PASS. Dry-run:cd deployment && ansible-playbook playbooks/site.yml --check --tags secretsconverges onmx99.jj commit -m "feat: per-cell postmaster password secret plumbing" config.go config_test.go deployment/roles/sovrnd/templates/sovrn.toml.j2 deployment/roles/zds/tasks/secret.yml deployment/scripts/provision-zds-secretsTask 4: Postmaster account + mapping record + saga wiring
Files: - Create:
internal/postmaster/postmaster.go(invite → createAccount → createSession → putRecord/deleteRecord against the cell ZDS loopback) - Modify:internal/appview/provision.go(ProvisionDomainbecomes the saga with compensations) - Test:internal/postmaster/postmaster_test.go, extendinternal/domain/activate_test.go-style provision testsNew
ProvisionDomainorder (each step idempotent; failure runs compensations in reverse, best-effort log-and-continue, original error returned): 1.domain.EnsureInactive(+ tenant precheck, unchanged). 2.dkim.Generate(now idempotent, Task 2). 3.store.CreateDomain(status: verifying)with upfront TID (Task 2). 4.Supervisor.Activateper-domain ZDS instance (existing rollback inside; burns slug/port tombstone — accepted). 5. Postmaster: mint invite code with admin token →com.atproto.server.createAccount{handle: postmaster.at.<domain>, email: postmaster@<domain>, password: <shared file>, inviteCode}→ onInvalidRequest/account-exists,createSessionwith shared password and verify DID matches expectation →putRecord(at.sovrn.domain.route/self)with{domain, mxHost: sovrn_mail_hostname, cell: <cell hostname>, status: active}. 6. Existing readiness gate →domain.Enable→ status active.Compensations:
deleteRecord(route/self)→ deactivate postmaster (com.atproto.admin.updateSubjectStatus, best-effort; handle abandoned) →Supervisor.Retire→store.DeleteDomain→dkim.DeleteForDomain→domain.Delete. Terminal errors (validation/conflict) → 4xx after rollback; transient (transport/timeout) → 5xx, caller retries andEnsure*converges.putRecordbody; per-step fault injection (fail at steps 3/4/5/6) asserts the exact compensation sequence ran and the original error surfaced.go test ./internal/postmaster/ ./internal/appview/ -v— Expected: FAIL.internal/postmasterclient + saga rewrite ofProvisionDomain.go test ./internal/postmaster/ ./internal/appview/ ./internal/domain/ ./internal/dkim/ -v— Expected: PASS.jj commit -m "feat: domain setup saga with postmaster route record" internal/postmaster internal/appviewTask 5: Retire path (route withdrawal before teardown)
Files: - Modify:
internal/appview/retire/teardown handler (whichever owns domain retire — checkdomain_list.gosiblings),internal/postmaster/postmaster.go(status-update helper)Order:
putRecord(status: retired)→ verify backup observed withdrawal (best-effort: short grace, documented) → existing teardown (accounts →dkim.DeleteForDomain→domain.Delete,Supervisor.Retire,store.DeleteDomain).putRecord{status:retired}precedes anyx:Domain/set destroy/Retirecall on the fakes.go test ./internal/appview/ -run TestRetire -v— Expected: FAIL.go test ./internal/appview/ -v— Expected: PASS.jj commit -m "feat: withdraw route record before domain teardown" internal/appview internal/postmasterTask 6: Orphan sweep (stateless, no journal table)
Files: - Modify: background sweep — extend the verifier/sweep surface (
internal/verifier/) or the reconciler sweep wherever domain lifecycle sweeps live; followListDomainsByStatuspattern (internal/store/store.go:66)Each cycle reports + repairs: Stalwart
x:Domain/queryIDs with nomail_domainsrow (delete after grace, audit-logged); DB rows stuckverifyingpast reap window with no route record (retry saga tail or reap per existingreapaftersemantics);retiringZDS registry rows past TTL. Metrics + audit entries per repair; never delete anything created within the grace window.go test ./internal/verifier/ -v— Expected: FAIL.go test ./internal/verifier/ -v— Expected: PASS.jj commit -m "feat: orphan sweep for partial domain setups" internal/verifier internal/storeTask 7: Backup box — inventory, Stalwart relay config, DNS, monitoring
Files: - Modify:
deployment/inventory/hosts.yml(newmail_backupgroup, e.g.backup.eu.sovrn.at),deployment/inventory/group_vars/all/sovrn.yml(backup hostname vars), per-hosthost_vars- Create:deployment/roles/stalwart_backup/(tasks + templates:MtaRoute Relayper cell pinned at cell hostname — neverMx;MtaOutboundStrategy.routebranching onrcpt_domain;MtaDeliveryScheduleTTL 5–7d with backoff;MtaStageRcpt.allowRelayingexception for backed-up domains;:25listener; ACME cert for backup hostname; firewall; vmagent/Vector edge wiring mirroringtelemetry_edge) - Modify:router.go:148(MailDefaults.MXgains pri-20),lexicons/at/sovrn/domain/defs.json(dnsState.mxdocuments both),internal/dnsprober+internal/verifierreadiness gate (accept pri-10 + pri-20), setup-guide UI copy (second MX line), hosted-zone wildcard automation - Modify:docs/runbooks/relay-drain-verify.md(queue-depth source = backup Stalwart queue listing),docs/deployment.mdmail-flow section (mark implemented)Key invariants (ADR-0009 D28 + Stalwart docs): route address is always the cell hostname, never
Mx(backup’s own hostname is in the MX set → loop); SPFmxalready covers the backup; DMARC unchanged. NoDomain/Accountobjects for backed-up domains on the backup (accept-all, no recipient validation — v1).just update backup.eu.sovrn.at(or cell equivalent) — Expected: green, Stalwart:25up, queue empty, no local domains.relay-drain-verify.mdchecks).jj commit -m "feat: pri-20 backup MX relay role and DNS" deployment/inventory deployment/roles/stalwart_backup router.go internal/dnsprober internal/verifier docs/Task 8: E2E verification on mx99 + docs close-out
mx99.eu.sovrn.at:domain.create(hosted) → assert saga artifacts (Stalwart domain, ZDS instance, postmaster account, route record, DB row) → kill primary path → send mail → assert backup queues → restore → assert drain + JMAP readability → retire domain → assertretiredwithdrawal precedes teardown.docs/deployment.md(mail flow implemented),docs/08-security-compliance.md(postmaster password in key inventory + rotation gap: no ZDS password-change API, compromise = rotate vault + recreate accounts), and this bug with results.docs/runbooks/recovery-primary-ip.md/relay-drain-verify.md.Explicit follow-ups (NOT this plan)
domain.createguard (read-only route-record check → 409 + redirect hint).Task 1 complete: at.sovrn.domain.route lexicon + validation (spec ✅, quality ✅ after fix loop: length/edge pins, CheckLabel dedup via exported pdslifecycle helper, trailing-dot consistency, UpdatedAt RFC3339, RouteRecord/ValidateRoute rename).
Task 2 complete: idempotent DKIM generate (query-first reuse, zero-Set on all branches) + upfront domain TID (spec ✅, quality ✅; timing-based TID test accepted as robust, deterministic-injection noted as optional follow-up).
Task 3 complete: per-cell postmaster password plumbing (spec ✅, quality ✅ + 3 nit fixes). Fix-forward from Task 1 also landed: at.sovrn.domain.route added to login scopes (full suite green).
Task 4 complete: domain setup saga (spec ✅, quality ✅ after fix loop: gofmt, shared HTTP client, detached rollback ctx, AdminToken gate, double-%w). Follow-up filed: 5e674f6 (teardown provenance — fix before supervision-on prod).
Task 5 complete: withdraw route record before teardown, wired in prod (spec ✅ after wiring fix, quality review pending on final diff).
Task 5 complete: withdraw route record before teardown, wired in prod (spec ✅, quality ✅ + 3 nit fixes).
Task 6 complete: stateless orphan sweep (spec ✅ incl. DNS-pending safety trace, quality ✅; Task 5 gap #3 closed — reap now Retires the instance).
Task 7 complete: backup relay role + DNS (spec ✅, quality ✅ after fix loop: day-2 reconverge via statefile, TTL int, nits). Follow-up filed: 0a0fb42 (jetstream route-sync watcher replacing manual domain-map var).
Live E2E checklist: backup-MX (bug d43eebe, Task 8) — OPERATOR-RUN ONLY
Agent performed docs/static checks only. No live hosts touched, no vault decrypted, no playbooks run. The operator executes every step below.
Preconditions
kkznturs f90c6c06(feat: pri-20 backup MX relay role and DNS) +nrwtqpxy b5888f15(fix: task 7 review findings — day-2 relay reconverge, TTL int, nits).just provision-zds-secrets <cell>(provides per-cellsovrn_postmaster_passwordalongside the fourzds_*keys)..ssh/configmaps its hostname.A. Bootstrap asserts (Task 7 Step 1)
mail_backupindeployment/inventory/hosts.yml(sovrn_backup_cells/sovrn_backup_domain_routes/sovrn_mail_hostname), per the group comment.mode=bootstrap), thenjust persist-bootstrap-secrets <backup-host>, review, commit.:25-only listeners (no submission/IMAP user ports); oneRelayroute per cell;remoteschedule TTL =sovrn_backup_queue_ttl_days(default 7d); ACME cert for the backup hostname; queue empty; no local Domain/Account objects for backed-up domains.B. Cell update asserts (Task 7 DNS half)
sovrn_backup_mxfleet-wide to the backup FQDN; converge cells.verifying).C. Negative tests
allowRelayingelse = false).Relayroute address names the cell hostname — never plainMx(backup hostname is in the MX set → loop).D. Fence-cell queue-hold + drain, retire ordering (runbook §§1–2)
putRecord{status:retired}precedes teardown (retired-withdrawal ordering), then confirm the queue TTL + orphan sweep bound the tail.--tags stalwart-backup-bootstrapreconverge on every onboard/retire).Record deltas here
Reply on this bug with: PASS/FAIL per group, measured queue depths and drain latency, and any rehearsal deltas for
relay-drain-verify.md(§4) orrecovery-primary-ip.md.Final review caught one blocker (Task 6 sweep unwired in prod — same class as Task 5 gap); fixed + reviewed: orphanSweepDeps supervision-gated over held registry, WithOrphanSweep chained. Stack approved for operator E2E. Remaining live work is the operator checklist (comment d9453be); open follow-ups: 5e674f6 (teardown provenance), 0a0fb42 (route-sync watcher).
Status (2026-10-06): the Ansible relay this was built on was never live-tested (the operator E2E in d9453be didn’t run before the hosts were torn down), and Ansible is being retired (3beadb2). The relay is being rebuilt on NixOS in b45a76f, on the shared netcup box. This issue’s acceptance (mail queues for an unreachable cell and drains on return) is the same as b45a76f’s; close this when that passes on staging (plan Phase 9).
Update (2026-10-06): the NixOS relay (b45a76f) passes a VM test with this issue’s acceptance: mail for a cell’s domain queues while the cell is down and is delivered when it returns. Changes from the Ansible design: cells take relayed mail on a firewalled 2525 instead of 25 (see b45a76f), and route-sync uses its own restricted Stalwart account. Close this after the live check on staging (plan Phase 9).
Live queue-and-drain passed (fleet Phase 9, 2026-10-08): mx99 (NixOS staging cell, arm64) Stalwart stopped 03:51:02; Gmail -> [email protected] fell back to the pri-20 MX mxb.eu.sovrn.at (relay on infra.mymood.at) at 03:51:55, SPF/DKIM/DMARC pass; queued 03:52:53, first attempt to mx99:2525 refused (v4+v6), rescheduled +5 min, expiry 7 d; Stalwart back 03:53:42; relay retry 03:57:55 on 2525 with STARTTLS, 250 at 03:58:22, queue empty; mx99 ingested it into the mailbox (relay listener 2525) 03:58:51. route-sync had applied the domain from mx99’s /backup/domains (domains:1). backupMx = mxb.eu.sovrn.at is set in sovrn fleet.nix and deployed to mx99.
Closing (user, 2026-10-08): done. The pri-20 backup MX queues while a cell is unreachable and drains when it returns, verified live on mx99 (Phases 9 and 10); backupMx = mxb.eu.sovrn.at is published to every domain.