Inbound DANE: publish TLSA records for the cells' and the relay's MX hostnames
closedGoal
Let DANE-aware senders (Gmail, Microsoft 365, Postfix with DANE, …) authenticate our MX hosts and refuse to downgrade TLS when delivering to sovrn-hosted domains.
Current state (2026-10-08)
- sovrn.at is DNSSEC-signed (KSK algorithm 13, DS published in .at).
- No TLSA records:
_25._tcp.mx99.eu.sovrn.atand_25._tcp.mxb.eu.sovrn.atare NXDOMAIN. - Outbound DANE is already on: since d4988da, Stalwart resolves through a local validating unbound.
What’s needed
- TLSA records at
_25._tcp.<mx-host>for every cell and for mxb.eu.sovrn.at. They live in our zone, so hosted (customer) domains add nothing as long as their MX points at our hostnames. - It only applies to customer domains that are themselves DNSSEC-signed, with a DS record at their registrar. Customers opt in through their DNS provider; unsigned domains keep ordinary opportunistic TLS.
Open question to answer first: key rollover
The usual record is 3 1 1 (SHA-256 of the server’s SPKI). That only survives Let’s Encrypt renewals if:
- the private key is reused on renewal, or
- the next key’s TLSA record is published before the switch (current + next, one TTL ahead).
Check whether Stalwart 0.16’s ACME client reuses the key or can pre-generate the next one. Alternative: 2 1 1 on the Let’s Encrypt intermediates. That is fragile because LE rotates intermediates. A TLSA record that stops matching makes DANE senders defer all mail, which is worse than having no record.
Steps
- Answer the rollover question; choose 3 1 1 with key reuse, or a rollover procedure.
- Decide where TLSA records are managed: sovrnd’s DNS automation, or Stalwart’s DnsManagement, so new cells get them automatically.
- Publish on mx99 first; verify with
posttls-finger -t30 -T180 -c -L verbose,summary <domain>and https://dane.sys4.de / internet.nl against a signed test domain. - Monitor: alert when the TLSA record doesn’t match the served certificate, e.g. a check after each ACME renewal.
- Roll out to mxb and the other cells.
6 Comments
Findings (2026-10-08)
Rollover: answered
Stalwart 0.16’s
AcmeProviderhasreuseKey(crates/common/src/network/acme/renew.rs: on renewal it loads the private key of the existing certificate with the same SAN set). Our plans (nix/stalwart/cell.plan.json,relay.plan.json) don’t set it, so it’s false and every renewal makes a new key. With"reuseKey": true,3 1 1survives LE renewals unchanged. Remaining risk: if the reuse lookup misses (registry lost, SAN set changed), a fresh key is issued silently. Monitoring (step 4) stays necessary.DNS backend: Marque works for TLSA
sovrn.at’s zone is the PDS record
at://did:plc:zpqnooj2vuoju6ssulbpdwxp/at.marque.dns/sovrn.at(served by stratus.mqdns.de, cirrus.mqdns.at, nimbus.mqdns.app). The lexicon’srecordTypeknownValues don’t list TLSA, but tested: - Added_25._tcp.dane-testTLSA3 1 1 <mx99 SPKI sha256>withputRecord(+swapRecord) using the rtw.run app password (secrets: identities/rtw_run_app_password). - Served by all three NS within ~45 s, RRSIG with the zone key (alg 13, tag 10739); 1.1.1.1 and 8.8.8.8 return it with AD,delv“fully validated”. - Removed again; all three NXDOMAIN, rest of zone unchanged. - Value format:"<usage> <selector> <matching> <hex>", name relative to the apex.Gotchas for automation: - The record is the whole zone: read, change only
_25._tcp.*entries, write back withswapRecord= the read CID; retry on swap failure. - Marque rewrites the record itself (~30 s after a write: reordered records, new createdAt, same content), so the CID changes underneath a writer. Always re-read before writing. - Stalwart’s DnsManagement can’t target Marque; a writer would be ours (sovrnd or a host job).Proposed plan
reuseKey: truein cell and relay plans; deploy (one change).3 1 1for mx99 and mxb in the Marque record (by hand or by a small writer).Decision: no TLSA for the backup relay (mxb)
The backup relay fails open on purpose. DANE is per MX host (RFC 7672): with a validated denial of
_25._tcp.mxb.eu.sovrn.at, DANE senders authenticate mx99 strictly and deliver to mxb with opportunistic TLS. A TLSA mismatch on mxb would be a latent fault (almost nothing delivers there) that surfaces exactly when the primary is down._25._tcp.mxb.eu.sovrn.atis NXDOMAIN with AD at 1.1.1.1 and 8.8.8.8; delv “negative response, fully validated”. That validated denial is what keeps mxb failing open; a SERVFAIL there would make DANE senders defer to mxb too.Step 1 done (in sovrn, not deployed)
nix/stalwart/cell.plan.json:AcmeProvider.reuseKey: true.stalwart-cli applyupserts are patches (only the plan’s properties are sent), so the account key is untouched. stalwart-plan check passes. Takes effect at the next renewal: mx99 keeps today’s key (SPKI sha256 6d4d9619…0a5c6c).Steps 2-3 done: mx99 TLSA live (2026-10-09)
_25._tcp.mx99.eu.sovrn.at TLSA 3 1 1 6d4d9619…bd0a5c6cwith the newjust update-dnsin servers (plans fromjust check-dns, writes the Marque record with goat putRecord + swapRecord). Deployed beforehand: cells’ AcmeProvider reuseKey (sovrn 9834ad9, mx99).just check-dns: all green (Marque record, all three NS, 1.1.1.1 AD on the TLSA; mxb’s TLSA name a validated NXDOMAIN).openssl s_client -starttls smtp -dane_tlsa_domain mx99.eu.sovrn.at -dane_tlsa_rrdata …→ “DANE TLSA 3 1 1 matched the EE certificate at depth 0”, Verify return code 0. Negative control with a wrong hash → code 65 “no matching DANE TLSA records”.Remaining: step 4 (alert when the TLSA stops matching: e.g. check-dns on a timer, or a check on the metrics box), and new cells get their TLSA through
just update-dnswhen they join hosts.json.Step 4: DANE watch on the metrics box (built, awaiting deploy)
sovrn 23fc06d
nix/modules/dane-check.nix(sovrn.daneCheck), enabled by the metrics role when it has hosts; servers hive.nix gives it the inventory’s cells (hosts) and relays (failOpenHosts).Step 4 deployed (2026-10-09)
infra.mymood.at on dkdnnz1…; sovrn-dane-check.timer listed (every 5 min), manual run Result=success. VM has sovrn_dane_tlsa_match{host=mx99.eu.sovrn.at}=1, sovrn_dane_tlsa_absent{host=mxb.eu.sovrn.at}=1, probe_success 1 for both, last-run timestamp present. vmalert group dane: DaneTlsaMismatch, DaneBackupTlsaPresent, DaneCheckStale all inactive / health ok.
Status of the plan: 1 (rollover: reuseKey) done, 2 (Marque via just update-dns) done, 3 (published + verified) done, 4 (watch + alerts) done. 5 (other cells): mx99 is the only cell; a new cell gets its TLSA from
just update-dnsonce it’s in hosts.json, and the watch picks it up from hive.nix. mxb deliberately has none.Alert tested end to end; closing (2026-10-09)
Deliberate test: the mx99 TLSA was deleted from the Marque record at 11:21:03Z. The 11:25:08 check logged the mismatch (tlsa_match 0), vmalert went pending at 11:26 and fired DaneTlsaMismatch at 11:36 (for 10m); the email arrived ~16 min after the deletion (≤5 min to the next check + ~1 min evaluation + 10 min for). Kept as is: two consecutive failing checks before paging.
Note: the alert text says DANE senders defer mail, which is true for a wrong record; a missing one (validated NXDOMAIN) fails open instead. Left as is by decision.
Record restored with
just update-dns(authoritative NS serving it again). Inbound DANE is live on mx99; mxb fails open by design; new cells are covered by just update-dns + the watch.