mxb hands non-local mail (its own DMARC reports) to mx99:2525, which rejects it

closed
#b1f7229 opened by agent Oct 9

Seen while working e9d14c1, in the mx99.eu.sovrn.at Stalwart journal on 2026-10-09:

04:25:53Z SMTP EHLO listenerId=relay localPort=2525 remoteIp=152.53.243.113 domain=mxb.eu.sovrn.at
04:25:53Z SMTP MAIL FROM [email protected]
04:25:53Z Relay not allowed (smtp.relay-not-allowed) [email protected]

mxb (the backup relay) tried to send its own DMARC aggregate report for google.com through mx99’s 2525 relay listener. mx99 is right to refuse: 2525 takes relayed mail only for the cell’s own domains.

To find out

  • Why mxb routes outbound mail that isn’t for a hosted domain to a cell at all. Check its outbound routing/strategy and whether every non-local destination falls through to the cell route.
  • Whether mxb should send DMARC/TLS reports at all, and if it should, through which path (SMTP2GO? direct?).
  • What happens to the rejected report on mxb: does it retry, bounce or pile up in its queue? Check mxb’s queue for more of these.

4 Comments

agent be1efe7 Oct 9

Root cause (2026-10-09)

internal/routesync/express.go strategyExpr builds mxb’s MtaOutboundStrategy.route as match rcpt_domain == '<served domain>' -> 'cell-<host>', else 'cell-<Cells[0]>'. Every message mxb generates for a domain it doesn’t serve therefore goes to the first cell’s 2525 ingress. The cell is right to refuse it (allowRelaying is limited to its own domains).

mxb journal, last 7 days: 1 queue.report-queued (DMARC aggregate for google.com, from [email protected]). Delivered to mx99, which answered 550 Relay not allowed, giving dsn-perm-fail, and the report was dropped. Nothing else is stuck in the queue.

What else falls into the same else branch

  • Aggregate/failure reports (DMARC, TLS-RPT, etc.) that mxb produces for mail it accepts on 25. Port-25 defaults turn these on.
  • DSNs to external senders. If mxb gives up on a cell after TTLDays, or sends a delay notice, the DSN to the original (external) sender takes the else route and is refused the same way. The bounce is lost and the sender is never told. No case so far, but this is the important one.

Facts for choosing a fix

  • mxb (152.53.243.113, the netcup box) can reach port 25 outbound: gmail-smtp-in answered with a 220. Its PTR is mxb.eu.sovrn.at.
  • mxb.eu.sovrn.at has no SPF TXT record.

Options

  1. Turn off report generation on the relay (plan.json). The cell generates reports for the mail it finally receives. This fixes the reports only.
  2. else 'mx' (direct delivery) for everything that isn’t a served domain. DSNs and any reports go out directly from mxb. Needs SPF for mxb.eu.sovrn.at (v=spf1 a -all; DSNs use the null sender, so the HELO identity is what gets checked). It also opens a direct outbound path from the relay; the RCPT-stage relaying expression stays closed, so it can’t become an open relay.
  3. else through SMTP2GO: the relay would need its own SMTP2GO credentials and a verified sender domain. Heavier.

Recommendation: 1 + 2. Reports off (nothing on the relay needs them), and else 'mx' so bounces reach the people sending to sovrn users.

agent b316fe7 Oct 9

Fix (chosen: reports off + direct MX), 2026-10-09

  • sovrn vkwxykot: strategyExpr else is now 'mx' (Stalwart’s built-in direct route). The relay plan sets sendFrequency/aggregateSendFrequency/failureSendFrequency to {"else":"disable"} on DmarcReportSettings, DkimReportSettings, SpfReportSettings and TlsReportSettings (disable makes Stalwart’s rate/frequency evaluation yield nothing). Checked: go test ./internal/routesync, flake checks stalwart-plan (real apply, idempotent) and relay (VM).
  • servers nunxsuvv: check-dns expects mxb.eu.sovrn.at TXT v=spf1 a -all and normalises dig’s quoted TXT. DRY=1 just update-dns plans exactly that one addition.
  • servers swlprpqo: sovrn input bumped. just diff-host infra.mymood.at: stalwart-converge, sovrn-route-sync unit, Stalwart restart trigger.

To roll out (user): 1) just update-dns (SPF first), 2) just deploy infra.mymood.at. Verify on mxb: the strategy’s else reads ‘mx’, and report settings read disable.

agent b11ff47 Oct 10

Verified live (2026-10-10)

  • DNS: mxb.eu.sovrn.at TXT "v=spf1 a -all" on all three nameservers.
  • mxb runs the diffed system (9wjn1z53…); route-sync applied the map (1 cell, 1 domain).
  • Live config read back as admin: strategy match test.kilimanjaro.io -> 'cell-mx99-eu-sovrn-at', else 'mx'; allowRelaying only test.kilimanjaro.io, else false; DMARC (aggregate, failure), DKIM, SPF and TLS report frequencies all disable.
  • End-to-end DSN: a message from [email protected] to [email protected] was injected on mxb:25 (queue 334112758932963328) and relayed to mx99, which answered 550 5.1.2. mxb queued a DSN (<> -> [email protected], queue 334112776347714048) and delivered it directly to in1-smtp.messagingengine.com over STARTTLS: 250 queued as C991611200AD.
  • Open-relay check: [email protected] -> [email protected] on mxb:25 refused with 550; nothing queued.

Closed.

agent bf1ef17 Oct 10

Bounce as received at Fastmail (2026-10-10)

The test DSN reached [email protected] and was filed as spam. The reason is backscatter detection, not reputation: - X-Backscatter: Yes, X-ME-VSCategory: bounce. The attached original was injected on mxb (HELO tmp.sb7lIbL8zn) and never passed through Fastmail’s outbound servers, so Fastmail treated it as a bounce for mail the user didn’t send. A DSN for a message really sent from Fastmail would carry Fastmail’s own headers and DKIM. - The sender side is clean: X-Spam-score: 0.7, sender reputation 500 (neutral), spf=pass smtp.helo=mxb.eu.sovrn.at (the new record), iprev/PTR pass, TLS 1.3. - Known minor gap, accepted: DSNs from mxb aren’t DKIM-signed (From: [email protected]). Revisit only if real bounces start landing in spam.