Stalwart inbound spam filter takes 25-60 s per message (likely DNS list lookups timing out)
closedObserved (2026-10-08, fleet Phase 9)
Every inbound message waits a long time between Stalwart’s last SMTP-session check (DMARC) and being queued, which is where the spam filter runs: - mx99 (cell, Hetzner, port 25, Gmail): DMARC 03:28:02 -> queued 03:28:31 (29 s) - relay mxb.eu.sovrn.at (infra.mymood.at, netcup, port 25, Gmail): DMARC 03:51:57 -> queued 03:52:53 (56 s) - mx99 relay listener (2525, from the relay): RCPT 03:58:25 -> queued 03:58:51 (26 s) Locally submitted/direct test mail (swaks from the shared box) showed the same shape. Delivery succeeds; the latency is all inside the session, so the sending MTA holds the connection open the whole time (risk: sender-side timeouts under load, slow fallback to the backup MX, more concurrent sessions).
Likely cause (to verify)
DNS-based lookups in Stalwart’s spam filter (DNSBL/URIBL/RBL checks, possibly pyzor or other network classifiers) timing out or waiting on slow resolvers: - Cells use Hetzner’s resolvers via systemd-resolved; the shared box uses netcup’s. Plain port 53 to public resolvers is blocked from some networks we use, and many public DNSBLs refuse queries from large public/cloud resolvers or rate-limit them (Spamhaus returns 127.255.255.x ‘public resolver’ codes). - Stalwart also logs that the resolver can’t validate DNSSEC (DANE disabled). - The relay VM test pins public.pyzor.org to 127.0.0.2 to fail fast, a hint that network classifiers block when unreachable.
To do
- Turn on debug/trace for the spam-filter stage on mx99 for one message and see which lookups take the time.
- Decide per check: keep with a short timeout, point at a resolver that DNSBLs accept (a local caching resolver on each host, e.g. unbound, querying the roots itself; or a Spamhaus DQS key), or disable (pyzor, URIBLs that time out).
- Set explicit, short timeouts for network checks in the Stalwart plan (nix/stalwart/*.plan.json) so the session never waits tens of seconds.
- Measure again on mx99 and the relay; target a few hundred ms.
- Consider DNSSEC validation in the same change (local validating resolver), which also re-enables DANE for outbound.
4 Comments
Diagnosis (2026-10-08, mx99, Stalwart 0.16.23)
Reproduced with a probe from infra.mymood.at directly to mx99:25: DMARC done 07:46:58.75, queued 07:47:30.07 (31 s). I captured packets on mx99 during the session and paired Stalwart’s queries to systemd-resolved (127.0.0.53) with the answers.
Cause: systemd-resolved’s
DNSSEC=allow-downgrade(serversmodules/base.nix), not the DNSBLs themselves. When an unsigned zone returns NXDOMAIN, resolved tries to prove the answer is insecure by sending DS queries down every label of the name. A reverse IPv6 DNSBL name has 32 nibble labels, and Hetzner’s resolvers SERVFAIL those DS queries after 1-1.6 s each (the list operators’ nameservers mishandle DS):--validate=noHow the 31 s splits: - rep.mailspike.net ~15 s (Stalwart’s resolver retried 9 times, then gave up at 15 s; resolved answered at 17.8 s) - bl.blocklist.de ~9 s - ~6 s for about 20 other DNSBL/URIBL lookups at 0.1-1 s each. Stalwart runs these one after another, so the time adds up.
The relay (infra.mymood.at, netcup resolvers) has the same resolved config, which fits its 56 s.
Second finding: several lists refuse our resolvers, so those checks currently do nothing. zen/dbl.spamhaus.org returns 127.255.255.254 (public resolver), list.dnswl.org returns 127.0.0.255, multi.uribl.com returns 127.0.0.1, and multi.surbl.org returns 127.0.0.254 (blocked/refused codes). pyzor’s TXT lookup is fast (73 ms) and isn’t involved.
Fix options
DNSSEC = falsefor resolved (fleet-wide in base.nix, or only on mail hosts). This removes the ~24 s of DS chasing. allow-downgrade gives little protection anyway.Fix deployed to mx99 (sovrn 17c7b74)
Stalwart now resolves through a local unbound (127.0.0.1:5335, UDP and TCP), which recurses from the root servers itself. The host’s systemd-resolved is unchanged.
Same probe as before (infra.mymood.at -> mx99:25), DMARC -> queued: - before: 31 s - after, unbound’s cache empty: 6.8 s - after, warm cache: 127 ms (whole SMTP session 0.6 s)
Spamhaus (127.0.0.10), dnswl (127.0.10.0) and URIBL (127.0.0.14) now give their real test-point answers instead of “refused”. SURBL still returns 127.0.0.254; check against SURBL’s own test entry. Normal-mode Stalwart no longer warns that the resolver can’t validate DNSSEC, so DANE is on for outbound. The warning only appears in the recovery-mode instance, before the plan switches the resolver.
Remaining: - Deploy to the relay (infra.mymood.at). - Cold-cache lookups still run one after another (6.8 s). Consider trimming the DNSBL list or shorter per-check timeouts. - Inbound DANE (TLSA records for the cells and mxb) is a separate follow-up.
Relay: unbound got no answers after deploy
On infra.mymood.at the new unbound answered nothing: netcup’s firewall policy only allows listed inbound ports, so replies to outbound UDP DNS (to random high ports) are dropped. Only netcup’s own resolvers get through. TCP DNS works over IPv4 and IPv6. Normal-mode Stalwart there logged “resolver cannot validate DNSSEC”. No mail reached the relay in that window.
Fix: servers 0e9be99 sets
tcp-upstream: yesfor unbound on netcup hosts (modules/netcup.nix). Stalwart on the relay needs a restart after that deploy, because it checks DNSSEC only at startup.Relay fixed (netcup firewall rule) and measured
netcup’s firewall policy now allows inbound UDP from source ports 53 (DNS), 123 (NTP) and 24441 (pyzor). With that in place, servers 41715a6 reverted the tcp-upstream workaround, so the relay’s unbound queries over plain UDP. The same block had also kept the relay’s clock from ever syncing; it was 31 s behind and has now been stepped and synced.
Pyzor had cost exactly one 5 s timeout per message on the relay. Its replies come from UDP port 24441, which the capture showed never arriving.
DMARC -> queued, warm cache:
End to end through the relay (relay:25 -> mx99:2525 -> inbox) takes about 2 s. Both Stalwarts run with DNSSEC validation, so DANE is on for outbound.
Leftovers, not blocking: the first lookups after a restart (empty cache) still run one after another (6.8 s on mx99, 9 s on the relay). Inbound DANE is bug 209274a.