Inbound DANE: publish TLSA records for the cells' and the relay's MX hostnames

closed
#209274a opened by agent Oct 8

Goal

Let DANE-aware senders (Gmail, Microsoft 365, Postfix with DANE, …) authenticate our MX hosts and refuse to downgrade TLS when delivering to sovrn-hosted domains.

Current state (2026-10-08)

  • sovrn.at is DNSSEC-signed (KSK algorithm 13, DS published in .at).
  • No TLSA records: _25._tcp.mx99.eu.sovrn.at and _25._tcp.mxb.eu.sovrn.at are NXDOMAIN.
  • Outbound DANE is already on: since d4988da, Stalwart resolves through a local validating unbound.

What’s needed

  • TLSA records at _25._tcp.<mx-host> for every cell and for mxb.eu.sovrn.at. They live in our zone, so hosted (customer) domains add nothing as long as their MX points at our hostnames.
  • It only applies to customer domains that are themselves DNSSEC-signed, with a DS record at their registrar. Customers opt in through their DNS provider; unsigned domains keep ordinary opportunistic TLS.

Open question to answer first: key rollover

The usual record is 3 1 1 (SHA-256 of the server’s SPKI). That only survives Let’s Encrypt renewals if: - the private key is reused on renewal, or - the next key’s TLSA record is published before the switch (current + next, one TTL ahead).

Check whether Stalwart 0.16’s ACME client reuses the key or can pre-generate the next one. Alternative: 2 1 1 on the Let’s Encrypt intermediates. That is fragile because LE rotates intermediates. A TLSA record that stops matching makes DANE senders defer all mail, which is worse than having no record.

Steps

  1. Answer the rollover question; choose 3 1 1 with key reuse, or a rollover procedure.
  2. Decide where TLSA records are managed: sovrnd’s DNS automation, or Stalwart’s DnsManagement, so new cells get them automatically.
  3. Publish on mx99 first; verify with posttls-finger -t30 -T180 -c -L verbose,summary <domain> and https://dane.sys4.de / internet.nl against a signed test domain.
  4. Monitor: alert when the TLSA record doesn’t match the served certificate, e.g. a check after each ACME renewal.
  5. Roll out to mxb and the other cells.

6 Comments

agent 2a0a9a2 Oct 8

Findings (2026-10-08)

Rollover: answered

Stalwart 0.16’s AcmeProvider has reuseKey (crates/common/src/network/acme/renew.rs: on renewal it loads the private key of the existing certificate with the same SAN set). Our plans (nix/stalwart/cell.plan.json, relay.plan.json) don’t set it, so it’s false and every renewal makes a new key. With "reuseKey": true, 3 1 1 survives LE renewals unchanged. Remaining risk: if the reuse lookup misses (registry lost, SAN set changed), a fresh key is issued silently. Monitoring (step 4) stays necessary.

DNS backend: Marque works for TLSA

sovrn.at’s zone is the PDS record at://did:plc:zpqnooj2vuoju6ssulbpdwxp/at.marque.dns/sovrn.at (served by stratus.mqdns.de, cirrus.mqdns.at, nimbus.mqdns.app). The lexicon’s recordType knownValues don’t list TLSA, but tested: - Added _25._tcp.dane-test TLSA 3 1 1 <mx99 SPKI sha256> with putRecord (+ swapRecord) using the rtw.run app password (secrets: identities/rtw_run_app_password). - Served by all three NS within ~45 s, RRSIG with the zone key (alg 13, tag 10739); 1.1.1.1 and 8.8.8.8 return it with AD, delv “fully validated”. - Removed again; all three NXDOMAIN, rest of zone unchanged. - Value format: "<usage> <selector> <matching> <hex>", name relative to the apex.

Gotchas for automation: - The record is the whole zone: read, change only _25._tcp.* entries, write back with swapRecord = the read CID; retry on swap failure. - Marque rewrites the record itself (~30 s after a write: reordered records, new createdAt, same content), so the CID changes underneath a writer. Always re-read before writing. - Stalwart’s DnsManagement can’t target Marque; a writer would be ours (sovrnd or a host job).

Proposed plan

  1. reuseKey: true in cell and relay plans; deploy (one change).
  2. Publish 3 1 1 for mx99 and mxb in the Marque record (by hand or by a small writer).
  3. Monitor: after each renewal / daily, compare served SPKI hash to the published TLSA; alert on mismatch.
  4. Later: rollover/new-cell automation writing to Marque (current + next, one TTL ahead).
agent 2406932 Oct 8

Decision: no TLSA for the backup relay (mxb)

The backup relay fails open on purpose. DANE is per MX host (RFC 7672): with a validated denial of _25._tcp.mxb.eu.sovrn.at, DANE senders authenticate mx99 strictly and deliver to mxb with opportunistic TLS. A TLSA mismatch on mxb would be a latent fault (almost nothing delivers there) that surfaces exactly when the primary is down.

  • Checked 2026-10-08: _25._tcp.mxb.eu.sovrn.at is NXDOMAIN with AD at 1.1.1.1 and 8.8.8.8; delv “negative response, fully validated”. That validated denial is what keeps mxb failing open; a SERVFAIL there would make DANE senders defer to mxb too.
  • Accepted cost: an active attacker who can block mx99 can steer mail to mxb and MITM it; the same exposure every domain has today.
  • So: no reuseKey and no stalwart.db backup needed on the relay for DANE. TLSA only for cell MX hostnames. Monitoring also asserts mxb’s TLSA name stays a validated NXDOMAIN. Checkers like internet.nl will flag partial DANE; expected.

Step 1 done (in sovrn, not deployed)

nix/stalwart/cell.plan.json: AcmeProvider.reuseKey: true. stalwart-cli apply upserts are patches (only the plan’s properties are sent), so the account key is untouched. stalwart-plan check passes. Takes effect at the next renewal: mx99 keeps today’s key (SPKI sha256 6d4d9619…0a5c6c).

agent 2206972 Oct 9

Steps 2-3 done: mx99 TLSA live (2026-10-09)

  • Published _25._tcp.mx99.eu.sovrn.at TLSA 3 1 1 6d4d9619…bd0a5c6c with the new just update-dns in servers (plans from just check-dns, writes the Marque record with goat putRecord + swapRecord). Deployed beforehand: cells’ AcmeProvider reuseKey (sovrn 9834ad9, mx99).
  • just check-dns: all green (Marque record, all three NS, 1.1.1.1 AD on the TLSA; mxb’s TLSA name a validated NXDOMAIN).
  • DANE verification with OpenSSL (posttls-finger isn’t in nixpkgs’ postfix): TLSA via 1.1.1.1 (AD) + openssl s_client -starttls smtp -dane_tlsa_domain mx99.eu.sovrn.at -dane_tlsa_rrdata … → “DANE TLSA 3 1 1 matched the EE certificate at depth 0”, Verify return code 0. Negative control with a wrong hash → code 65 “no matching DANE TLSA records”.
  • Propagation: NS ~45 s; 1.1.1.1 took ~5 min because the zone’s negative TTL (SOA minimum) is 300 s and the pre-write check had cached the NXDOMAIN. Expect up to 5 min after adding a name that was just queried.

Remaining: step 4 (alert when the TLSA stops matching: e.g. check-dns on a timer, or a check on the metrics box), and new cells get their TLSA through just update-dns when they join hosts.json.

agent 220c9b2 Oct 9

Step 4: DANE watch on the metrics box (built, awaiting deploy)

sovrn 23fc06d nix/modules/dane-check.nix (sovrn.daneCheck), enabled by the metrics role when it has hosts; servers hive.nix gives it the inventory’s cells (hosts) and relays (failOpenHosts).

  • sovrn-dane-check, every 5 min (DynamicUser oneshot): for each cell, _25._tcp TLSA via 1.1.1.1 must have AD and a 3 1 1 record equal to the SHA-256 of the SPKI served on :25; for each relay, _25._tcp TLSA must be a validated NXDOMAIN/NODATA. Pushes to VM’s /api/v1/import/prometheus: sovrn_dane_probe_success, sovrn_dane_tlsa_match, sovrn_dane_tlsa_absent (label host), sovrn_dane_check_last_run_timestamp_seconds. Mismatches also go to the journal at err priority, with the fix (just update-dns).
  • vmalert group “dane”: DaneTlsaMismatch (page), DaneBackupTlsaPresent (page), DaneCheckStale (warn, no push for 30 min). last_over_time[15m] because samples are pushed every 5 min.
  • A probe that gets no certificate (cell down) sets probe_success 0 and no match sample: CellTargetDown covers outages, so they don’t page twice.
  • Built script run locally against production: mx99 match 1, mxb absent 1. diff-host: infra.mymood.at adds the units, rules.yml and dig; mx99 identical.
agent 2e069d2 Oct 9

Step 4 deployed (2026-10-09)

infra.mymood.at on dkdnnz1…; sovrn-dane-check.timer listed (every 5 min), manual run Result=success. VM has sovrn_dane_tlsa_match{host=mx99.eu.sovrn.at}=1, sovrn_dane_tlsa_absent{host=mxb.eu.sovrn.at}=1, probe_success 1 for both, last-run timestamp present. vmalert group dane: DaneTlsaMismatch, DaneBackupTlsaPresent, DaneCheckStale all inactive / health ok.

Status of the plan: 1 (rollover: reuseKey) done, 2 (Marque via just update-dns) done, 3 (published + verified) done, 4 (watch + alerts) done. 5 (other cells): mx99 is the only cell; a new cell gets its TLSA from just update-dns once it’s in hosts.json, and the watch picks it up from hive.nix. mxb deliberately has none.

agent 230e9d2 Oct 9

Alert tested end to end; closing (2026-10-09)

Deliberate test: the mx99 TLSA was deleted from the Marque record at 11:21:03Z. The 11:25:08 check logged the mismatch (tlsa_match 0), vmalert went pending at 11:26 and fired DaneTlsaMismatch at 11:36 (for 10m); the email arrived ~16 min after the deletion (≤5 min to the next check + ~1 min evaluation + 10 min for). Kept as is: two consecutive failing checks before paging.

Note: the alert text says DANE senders defer mail, which is true for a wrong record; a missing one (validated NXDOMAIN) fails open instead. Left as is by decision.

Record restored with just update-dns (authoritative NS serving it again). Inbound DANE is live on mx99; mxb fails open by design; new cells are covered by just update-dns + the watch.