Backup MX (pri-20): store-and-forward relay with postmaster route records

closed
#d43eebe opened by agent Sep 18

Goal

Stand up a single pri-20 backup MX: a Stalwart store-and-forward relay that queues inbound mail while a cell is unreachable and drains when it returns (ADR-0009 D28). No recipient validation in v1.

Decisions locked in discussion (2026-09-18)

  • Routing source (option b): new postmaster-authored public record at.sovrn.domain.route in the per-domain postmaster’s own repo (postmaster.at.<domain>, ADR-0009 D32), key: literal:self. Fields: domain, mxHost, cell, status (active|retired, extensible), updatedAt. Backup jetstreams the collection, filters authorDid == postmaster DID, and verifies resolveHandle(postmaster.at.<domain>) == author plus domain-suffix match (spoof-proof, no cell coupling).
  • Rejected: overloading user-authored at.sovrn.mail.service (untrusted input, per-user cardinality, coverage gap before first login).
  • Saga: domain setup becomes atomic-ish — Stalwart EnsureInactive → DKIM (made idempotent) → DB row (verifying, TID minted upfront) → ZDS Supervisor.Activate → postmaster account (invite via admin token) → mapping putRecord last (externally-visible moment). Failure → reverse compensations, best-effort, original error preserved. Postmaster-account rollback is best-effort (deactivate + abandon handle; retry reuses handle via name-taken backstop).
  • Durability: synchronous compensations plus stateless orphan sweep (no new journal table initially): Stalwart domains without DB rows, stale verifying rows with no mapping record, retiring ZDS instances past TTL.
  • Postmaster credentials: one shared pre-provisioned password per cell (new vault var → 0600 file → [postmaster] passwordfile, same pattern as stalwart-secret). Verified against ZDS source: no admin impersonation exists (repo writes require the account’s own bearer; admin token gates only invites/takedown/sessions), and ZDS has no password-change endpoint, so per-domain random passwords would need unsustainable vault round-trips and still wouldn’t rotate. Envelope-encryption (option 2) deferred to D32 KMS-per-authority work.
  • Status field: domain-level status in the route record covers backup needs (active|retired; retire writes retired before teardown). Account-level enforcement (unpaid → block IMAP/JMAP, spam → block outgoing) stays Stalwart-side; record-status-vs-labeler is a separate spike, not this work.
  • Known v1 gap (accepted): cross-cell duplicate domain.create before first login is possible (per-cell DBs); login redirect + create-only service record resolve conflicts after first login. A read-only guard (check public route records in domain.create → 409 + redirect hint) is a fast follow.

Non-goals

Recipient validation on the backup, pri-30, per-domain passwords/KMS, the labeler spike, the cross-cell create guard (fast follow).

15 Comments

agent d84d33e Sep 18

Backup MX Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: A single pri-20 Stalwart relay queues mail for unreachable cells and drains on recovery, driven by postmaster-authored at.sovrn.domain.route records with atomic-ish domain setup.

Architecture: Cell setup runs a saga (Stalwart → DKIM → DB → ZDS → postmaster account → mapping record last) with reverse compensations plus an orphan sweep; the backup box is a new inventory group running Stalwart with per-cell Relay routes and a 5–7d queue, learning domain→cell mappings from the public jetstream.

Tech Stack: Go (sovrnd), Stalwart 0.16.x Registry API, ZDS XRPC, Ansible/systemd, jetstream firehose.


Task 1: at.sovrn.domain.route lexicon + types + validation

Files: - Create: lexicons/at/sovrn/domain/route.json - Modify: api/sovrn/ (regenerate via cmd/sovrn-lexgen — check Justfile/Makefile lexgen target) - Create: internal/postmaster/route.go (record type + validation) - Test: internal/postmaster/route_test.go

Record shape (key: literal:self, in postmaster’s public repo):

{
  "lexicon": 1,
  "id": "at.sovrn.domain.route",
  "defs": {
    "main": {
      "type": "record",
      "description": "Backup-MX routing claim, authored by the per-domain postmaster in its own public repo. One per repo.",
      "key": "literal:self",
      "record": {
        "type": "object",
        "required": ["domain", "mxHost", "cell", "status"],
        "properties": {
          "domain": { "type": "string", "maxLength": 253 },
          "mxHost": { "type": "string", "maxLength": 253, "description": "Cell MX hostname the backup relays to, e.g. mx1.eu.sovrn.at." },
          "cell": { "type": "string", "maxLength": 253, "description": "Owning cell hostname (Relay route target)." },
          "status": { "type": "string", "enum": ["active", "retired"], "description": "retired withdraws the route before teardown." },
          "updatedAt": { "type": "string", "format": "datetime" }
        }
      }
    }
  }
}

Validation rules in internal/postmaster/route.go (mirror internal/pdslifecycle/handles.go style — sentinel errors, table tests): - domain, mxHost, cell lowercased, non-empty, hostname-label rules via existing checkLabel-equivalent; domain must equal the postmaster handle suffix (postmaster.at.<domain>). - status must be active|retired; unknown values rejected (forward-compat: fail closed, matching cellmatch.go posture). - Author check helper: VerifyAuthor(ctx, domain, authorDID, resolveHandle) — resolves postmaster.at.<domain> and requires equality with authorDID.

  • [ ] Step 1: Write the failing test (internal/postmaster/route_test.go): table test for Validate (valid record passes; bad hostname, unknown status, domain/handle-suffix mismatch fail) and VerifyAuthor with a fake resolver (matching DID passes, mismatch fails).
  • [ ] Step 2: Run test to verify it fails — Run: go test ./internal/postmaster/ -run TestRoute -v — Expected: FAIL (package undefined).
  • [ ] Step 3: Write lexicon JSON + regenerate Go types + minimal route.go implementing the rules above.
  • [ ] Step 4: Run test to verify it passes — Run: go test ./internal/postmaster/ -v — Expected: PASS.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: at.sovrn.domain.route lexicon and validation" lexicons/at/sovrn/domain/route.json api/sovrn internal/postmaster

Task 2: Saga prerequisites — idempotent DKIM + upfront TID

Files: - Modify: internal/dkim/dkim.go (Generate gains query-first reuse) - Modify: internal/appview/provision.go (TID minted before Stalwart calls) - Test: internal/dkim/dkim_test.go, internal/appview/ provision tests

Today dkim.Generate mints a fresh DkimSignature per call — saga retry would stack signatures. Change: query DkimSignature by domainId first; if present, read back its public half and return without creating (same Key derivation as lines 73–96); else current create path.

TID: move id = store.NewID() in ProvisionDomain (provision.go:145) above the Stalwart Ensure call and thread it through, so crash-retry reuses the domain name as the idempotency key instead of orphaning under a new TID.

  • [ ] Step 1: Write failing tests — DKIM: stub client preloaded with an existing signature for the domain asserts zero x:DkimSignature/set calls and correct Key return. Provision: test asserting the DB row ID is fixed before/after a Stalwart retry (no duplicate CreateDomain IDs).
  • [ ] Step 2: Run to verify they fail — Run: go test ./internal/dkim/ ./internal/appview/ -v — Expected: FAIL.
  • [ ] Step 3: Implement query-first Generate + TID hoist.
  • [ ] Step 4: Run to verify pass — Run: go test ./internal/dkim/ ./internal/appview/ -v — Expected: PASS.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: idempotent DKIM generate and upfront domain TID" internal/dkim internal/appview

Task 3: Per-cell postmaster password secret plumbing

Files: - Modify: deployment/inventory/host_vars/<cell>/vault.yml pattern (document, operator step — never paste real values) - Create/modify Ansible task: extend deployment/roles/zds/tasks/secret.yml pattern with a postmaster-password file task, or a new task in deployment/roles/sovrn_secrets/tasks/main.yml - Modify: deployment/roles/sovrnd/templates/sovrn.toml.j2 ([postmaster] passwordfile), config.go (Postmaster struct + bindEnv + ValidateSecrets), deployment/scripts/provision-zds-secrets (generate 32-hex alongside existing keys)

New vault var sovrn_postmaster_password (per-cell, 32 hex like the Stalwart credential) → 0600 file /etc/sovrn/secrets/postmaster-password → sovrn.toml [postmaster] passwordfile. sovrnd refuses to start the domain-setup path when the file is missing (same posture as stalwart.secretfile, config.go:336).

  • [ ] Step 1: Write failing config test — config_test.go: ValidateSecrets without postmaster.passwordfile fails mentioning it (mirror the existing relay.apisecretfile test, config_test.go:60-82).
  • [ ] Step 2: Run to verify it fails — Run: go test . -run TestValidateSecrets -v — Expected: FAIL.
  • [ ] Step 3: Implement struct field + default + env binding + validation; Ansible task + toml template line; provision-zds-secrets key spec entry.
  • [ ] Step 4: Run to verify it passes — Run: go test . -v — Expected: PASS. Dry-run: cd deployment && ansible-playbook playbooks/site.yml --check --tags secrets converges on mx99.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: per-cell postmaster password secret plumbing" config.go config_test.go deployment/roles/sovrnd/templates/sovrn.toml.j2 deployment/roles/zds/tasks/secret.yml deployment/scripts/provision-zds-secrets

Task 4: Postmaster account + mapping record + saga wiring

Files: - Create: internal/postmaster/postmaster.go (invite → createAccount → createSession → putRecord/deleteRecord against the cell ZDS loopback) - Modify: internal/appview/provision.go (ProvisionDomain becomes the saga with compensations) - Test: internal/postmaster/postmaster_test.go, extend internal/domain/activate_test.go-style provision tests

New ProvisionDomain order (each step idempotent; failure runs compensations in reverse, best-effort log-and-continue, original error returned): 1. domain.EnsureInactive (+ tenant precheck, unchanged). 2. dkim.Generate (now idempotent, Task 2). 3. store.CreateDomain(status: verifying) with upfront TID (Task 2). 4. Supervisor.Activate per-domain ZDS instance (existing rollback inside; burns slug/port tombstone — accepted). 5. Postmaster: mint invite code with admin token → com.atproto.server.createAccount{handle: postmaster.at.<domain>, email: postmaster@<domain>, password: <shared file>, inviteCode} → on InvalidRequest/account-exists, createSession with shared password and verify DID matches expectation → putRecord(at.sovrn.domain.route/self) with {domain, mxHost: sovrn_mail_hostname, cell: <cell hostname>, status: active}. 6. Existing readiness gate → domain.Enable → status active.

Compensations: deleteRecord(route/self) → deactivate postmaster (com.atproto.admin.updateSubjectStatus, best-effort; handle abandoned) → Supervisor.Retire → store.DeleteDomain → dkim.DeleteForDomain → domain.Delete. Terminal errors (validation/conflict) → 4xx after rollback; transient (transport/timeout) → 5xx, caller retries and Ensure* converges.

  • [ ] Step 1: Write failing tests — fake ZDS server (httptest, loopback XRPC stubs for invite/createAccount/createSession/putRecord): happy path asserts call order and final putRecord body; per-step fault injection (fail at steps 3/4/5/6) asserts the exact compensation sequence ran and the original error surfaced.
  • [ ] Step 2: Run to verify they fail — Run: go test ./internal/postmaster/ ./internal/appview/ -v — Expected: FAIL.
  • [ ] Step 3: Implement internal/postmaster client + saga rewrite of ProvisionDomain.
  • [ ] Step 4: Run to verify they pass — Run: go test ./internal/postmaster/ ./internal/appview/ ./internal/domain/ ./internal/dkim/ -v — Expected: PASS.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: domain setup saga with postmaster route record" internal/postmaster internal/appview

Task 5: Retire path (route withdrawal before teardown)

Files: - Modify: internal/appview/ retire/teardown handler (whichever owns domain retire — check domain_list.go siblings), internal/postmaster/postmaster.go (status-update helper)

Order: putRecord(status: retired) → verify backup observed withdrawal (best-effort: short grace, documented) → existing teardown (accounts → dkim.DeleteForDomain → domain.Delete, Supervisor.Retire, store.DeleteDomain).

  • [ ] Step 1: Write failing test — retiring a domain asserts putRecord{status:retired} precedes any x:Domain/set destroy / Retire call on the fakes.
  • [ ] Step 2: Run to verify it fails — Run: go test ./internal/appview/ -run TestRetire -v — Expected: FAIL.
  • [ ] Step 3: Implement retire ordering.
  • [ ] Step 4: Run to verify it passes — Run: go test ./internal/appview/ -v — Expected: PASS.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: withdraw route record before domain teardown" internal/appview internal/postmaster

Task 6: Orphan sweep (stateless, no journal table)

Files: - Modify: background sweep — extend the verifier/sweep surface (internal/verifier/) or the reconciler sweep wherever domain lifecycle sweeps live; follow ListDomainsByStatus pattern (internal/store/store.go:66)

Each cycle reports + repairs: Stalwart x:Domain/query IDs with no mail_domains row (delete after grace, audit-logged); DB rows stuck verifying past reap window with no route record (retry saga tail or reap per existing reapafter semantics); retiring ZDS registry rows past TTL. Metrics + audit entries per repair; never delete anything created within the grace window.

  • [ ] Step 1: Write failing tests — seeded fakes (orphan Stalwart domain, stale verifying row, old retiring record) assert each is flagged/repaired and fresh rows are untouched.
  • [ ] Step 2: Run to verify they fail — Run: go test ./internal/verifier/ -v — Expected: FAIL.
  • [ ] Step 3: Implement sweep extension.
  • [ ] Step 4: Run to verify they pass — Run: go test ./internal/verifier/ -v — Expected: PASS.
  • [ ] Step 5: Commit — Run: jj commit -m "feat: orphan sweep for partial domain setups" internal/verifier internal/store

Task 7: Backup box — inventory, Stalwart relay config, DNS, monitoring

Files: - Modify: deployment/inventory/hosts.yml (new mail_backup group, e.g. backup.eu.sovrn.at), deployment/inventory/group_vars/all/sovrn.yml (backup hostname vars), per-host host_vars - Create: deployment/roles/stalwart_backup/ (tasks + templates: MtaRoute Relay per cell pinned at cell hostname — never Mx; MtaOutboundStrategy.route branching on rcpt_domain; MtaDeliverySchedule TTL 5–7d with backoff; MtaStageRcpt.allowRelaying exception for backed-up domains; :25 listener; ACME cert for backup hostname; firewall; vmagent/Vector edge wiring mirroring telemetry_edge) - Modify: router.go:148 (MailDefaults.MX gains pri-20), lexicons/at/sovrn/domain/defs.json (dnsState.mx documents both), internal/dnsprober + internal/verifier readiness gate (accept pri-10 + pri-20), setup-guide UI copy (second MX line), hosted-zone wildcard automation - Modify: docs/runbooks/relay-drain-verify.md (queue-depth source = backup Stalwart queue listing), docs/deployment.md mail-flow section (mark implemented)

Key invariants (ADR-0009 D28 + Stalwart docs): route address is always the cell hostname, never Mx (backup’s own hostname is in the MX set → loop); SPF mx already covers the backup; DMARC unchanged. No Domain/Account objects for backed-up domains on the backup (accept-all, no recipient validation — v1).

  • [ ] Step 1: Converge backup host to empty queue — Run: just update backup.eu.sovrn.at (or cell equivalent) — Expected: green, Stalwart :25 up, queue empty, no local domains.
  • [ ] Step 2: Live route test — inject a message for a cell domain via the backup with the cell firewalled off; assert queued, then drain on unblock with no loop signatures (relay-drain-verify.md checks).
  • [ ] Step 3: Commit — Run: jj commit -m "feat: pri-20 backup MX relay role and DNS" deployment/inventory deployment/roles/stalwart_backup router.go internal/dnsprober internal/verifier docs/

Task 8: E2E verification on mx99 + docs close-out

  • [ ] Rehearse full lifecycle on mx99.eu.sovrn.at: domain.create (hosted) → assert saga artifacts (Stalwart domain, ZDS instance, postmaster account, route record, DB row) → kill primary path → send mail → assert backup queues → restore → assert drain + JMAP readability → retire domain → assert retired withdrawal precedes teardown.
  • [ ] Update docs/deployment.md (mail flow implemented), docs/08-security-compliance.md (postmaster password in key inventory + rotation gap: no ZDS password-change API, compromise = rotate vault + recreate accounts), and this bug with results.
  • [ ] Record rehearsal deltas in docs/runbooks/recovery-primary-ip.md / relay-drain-verify.md.

Explicit follow-ups (NOT this plan)

  • Cross-cell domain.create guard (read-only route-record check → 409 + redirect hint).
  • Record-status-vs-labeler spike for account-level enforcement states.
  • KMS-per-postmaster migration (envelope encryption rides that work).
  • pri-30 / multi-region backup.
agent d44038e Sep 18

Task 1 complete: at.sovrn.domain.route lexicon + validation (spec ✅, quality ✅ after fix loop: length/edge pins, CheckLabel dedup via exported pdslifecycle helper, trailing-dot consistency, UpdatedAt RFC3339, RouteRecord/ValidateRoute rename).

agent d54e35e Sep 18

Task 2 complete: idempotent DKIM generate (query-first reuse, zero-Set on all branches) + upfront domain TID (spec ✅, quality ✅; timing-based TID test accepted as robust, deterministic-injection noted as optional follow-up).

agent dd4b3be Sep 18

Task 3 complete: per-cell postmaster password plumbing (spec ✅, quality ✅ + 3 nit fixes). Fix-forward from Task 1 also landed: at.sovrn.domain.route added to login scopes (full suite green).

agent dd4033e Sep 18

Task 4 complete: domain setup saga (spec ✅, quality ✅ after fix loop: gofmt, shared HTTP client, detached rollback ctx, AdminToken gate, double-%w). Follow-up filed: 5e674f6 (teardown provenance — fix before supervision-on prod).

agent d74733e Sep 18

Task 5 complete: withdraw route record before teardown, wired in prod (spec ✅ after wiring fix, quality review pending on final diff).

agent d04e33e Sep 18

Task 5 complete: withdraw route record before teardown, wired in prod (spec ✅, quality ✅ + 3 nit fixes).

agent de4f3be Sep 18

Task 6 complete: stateless orphan sweep (spec ✅ incl. DNS-pending safety trace, quality ✅; Task 5 gap #3 closed — reap now Retires the instance).

agent da4138e Sep 18

Task 7 complete: backup relay role + DNS (spec ✅, quality ✅ after fix loop: day-2 reconverge via statefile, TTL int, nits). Follow-up filed: 0a0fb42 (jetstream route-sync watcher replacing manual domain-map var).

agent d9453be Sep 18

Live E2E checklist: backup-MX (bug d43eebe, Task 8) — OPERATOR-RUN ONLY

Agent performed docs/static checks only. No live hosts touched, no vault decrypted, no playbooks run. The operator executes every step below.

Preconditions

  • [ ] Deployed code commits: kkznturs f90c6c06 (feat: pri-20 backup MX relay role and DNS) + nrwtqpxy b5888f15 (fix: task 7 review findings — day-2 relay reconverge, TTL int, nits).
  • [ ] Cell vault provisioned: just provision-zds-secrets <cell> (provides per-cell sovrn_postmaster_password alongside the four zds_* keys).
  • [ ] Backup box ordered and reachable over its primary IPv6; .ssh/config maps its hostname.

A. Bootstrap asserts (Task 7 Step 1)

  • [ ] Add backup host to mail_backup in deployment/inventory/hosts.yml
     + `deployment/inventory/host_vars/<backup-host>/vars.yml`
    
    (sovrn_backup_cells / sovrn_backup_domain_routes / sovrn_mail_hostname), per the group comment.
  • [ ] Bootstrap (mode=bootstrap), then just persist-bootstrap-secrets <backup-host>, review, commit.
  • [ ] Assert: :25-only listeners (no submission/IMAP user ports); one Relay route per cell; remote schedule TTL = sovrn_backup_queue_ttl_days (default 7d); ACME cert for the backup hostname; queue empty; no local Domain/Account objects for backed-up domains.

B. Cell update asserts (Task 7 DNS half)

  • [ ] Set sovrn_backup_mx fleet-wide to the backup FQDN; converge cells.
  • [ ] Assert: published MX = pri-10 cell + pri-20 backup; dnsprober / verifier readiness gate green (missing backup keeps domain verifying).

C. Negative tests

  • [ ] RCPT for an unknown (unmapped) domain is rejected by the backup (empty-map-converges-CLOSED: allowRelaying else = false).
  • [ ] Every Relay route address names the cell hostname — never plain Mx (backup hostname is in the MX set → loop).

D. Fence-cell queue-hold + drain, retire ordering (runbook §§1–2)

  • [ ] Fence the cell (recovery-primary-ip.md: stop services, power off); inject mail for a cell domain via the backup → assert queued (relay-drain-verify.md §1: depth + oldest-queued timestamp).
  • [ ] Restore the cell → assert drain + JMAP readability, no relay→cell→relay loop signatures, DSNs sane (relay-drain-verify.md §2).
  • [ ] Retire a domain → assert putRecord{status:retired} precedes teardown (retired-withdrawal ordering), then confirm the queue TTL + orphan sweep bound the tail.
  • [ ] Manual domain-map sync followed throughout (relay-drain-verify.md §4: map updated + --tags stalwart-backup-bootstrap reconverge on every onboard/retire).

Record deltas here

Reply on this bug with: PASS/FAIL per group, measured queue depths and drain latency, and any rehearsal deltas for relay-drain-verify.md (§4) or recovery-primary-ip.md.

agent dd483be Sep 18

Final review caught one blocker (Task 6 sweep unwired in prod — same class as Task 5 gap); fixed + reviewed: orphanSweepDeps supervision-gated over held registry, WithOrphanSweep chained. Stack approved for operator E2E. Remaining live work is the operator checklist (comment d9453be); open follow-ups: 5e674f6 (teardown provenance), 0a0fb42 (route-sync watcher).

agent d9453fe Oct 6

Status (2026-10-06): the Ansible relay this was built on was never live-tested (the operator E2E in d9453be didn’t run before the hosts were torn down), and Ansible is being retired (3beadb2). The relay is being rebuilt on NixOS in b45a76f, on the shared netcup box. This issue’s acceptance (mail queues for an unreachable cell and drains on return) is the same as b45a76f’s; close this when that passes on staging (plan Phase 9).

agent d04739e Oct 6

Update (2026-10-06): the NixOS relay (b45a76f) passes a VM test with this issue’s acceptance: mail for a cell’s domain queues while the cell is down and is delivered when it returns. Changes from the Ansible design: cells take relayed mail on a firewalled 2525 instead of 25 (see b45a76f), and route-sync uses its own restricted Stalwart account. Close this after the live check on staging (plan Phase 9).

agent d94133e Oct 8

Live queue-and-drain passed (fleet Phase 9, 2026-10-08): mx99 (NixOS staging cell, arm64) Stalwart stopped 03:51:02; Gmail -> [email protected] fell back to the pri-20 MX mxb.eu.sovrn.at (relay on infra.mymood.at) at 03:51:55, SPF/DKIM/DMARC pass; queued 03:52:53, first attempt to mx99:2525 refused (v4+v6), rescheduled +5 min, expiry 7 d; Stalwart back 03:53:42; relay retry 03:57:55 on 2525 with STARTTLS, 250 at 03:58:22, queue empty; mx99 ingested it into the mailbox (relay listener 2525) 03:58:51. route-sync had applied the domain from mx99’s /backup/domains (domains:1). backupMx = mxb.eu.sovrn.at is set in sovrn fleet.nix and deployed to mx99.

agent dc4a36e Oct 8

Closing (user, 2026-10-08): done. The pri-20 backup MX queues while a cell is unreachable and drains when it returns, verified live on mx99 (Phases 9 and 10); backupMx = mxb.eu.sovrn.at is published to every domain.