2 files changed,
+134,
-1
+2,
-1
1@@ -5,7 +5,7 @@ The fleet: every NixOS host, deployed with [Colmena](https://github.com/zhaofeng
2 | Host | Provider | Roles | Notes |
3 | --- | --- | --- | --- |
4 | `infra.rtw.run` | Hetzner | `forge` | Git hosting (soft-serve, kilimanjaro.io). Legacy layout (installed with nixos-infect), nixos-26.05 |
5-| `infra.mymood.at` | netcup | `moods`, `sovrn-metrics`, `sovrn-relay` | The shared box: moods, sovrn's metrics stack (`infra.sovrn.at`) and backup MX (`mxb.eu.sovrn.at`). Inbound ports also need netcup's firewall policy |
6+| `infra.mymood.at` | netcup | `moods`, `sovrn-metrics`, `sovrn-relay` | The shared box: moods, sovrn's metrics stack (`infra.sovrn.at`) and backup MX (`mxb.eu.sovrn.at`). Inbound ports and UDP replies also need netcup's firewall policy ([provider firewalls](docs/runbooks/provider-firewalls.md)) |
7 | `mx99.eu.sovrn.at` | Hetzner (CAX11, arm64) | `sovrn-cell` | sovrn's staging cell; also a Nix remote builder for aarch64 |
8
9 ## Layout
10@@ -19,6 +19,7 @@ The fleet: every NixOS host, deployed with [Colmena](https://github.com/zhaofeng
11 | `roles/` | Roles that live here (`forge`). Project roles come from the project flakes. |
12 | `hosts/<fqdn-with-dashes>/` | Legacy hosts only: hand-written hardware, networking, stateVersion. `hosts/common/` is shared by them. |
13 | `keys/admins.pub` | Root SSH keys for standard hosts. |
14+| `docs/runbooks/provider-firewalls.md` | The rules for provider firewalls (netcup's policy, Hetzner Cloud Firewalls) in front of the NixOS firewall, per host, with the reason for each. |
15
16 ## Workflow
17
+132,
-0
1@@ -0,0 +1,132 @@
2+# Provider firewalls
3+
4+Every host's packet filter is the NixOS firewall (`modules/base.nix`: on,
5+stateful, only SSH open; roles open their own ports). A provider firewall in
6+front of it (netcup's firewall policy, a Hetzner Cloud Firewall) is a second
7+layer that lives outside this repo. It must let through at least what the
8+host opens, and, if it doesn't track connections, the replies to what the
9+host sends. This page is the list, and the reason for each entry, so the
10+rules can be set by hand today and generated through the providers' APIs
11+later.
12+
13+**Convention:** one catch-all rule allows all outbound traffic; every other
14+rule is for inbound TCP or UDP. So the tables below are inbound only.
15+
16+The NixOS config is the source of truth for inbound ports. Check this page
17+against it after changing a role:
18+
19+ nix eval --json '.#nixosConfigurations."<fqdn>".config.networking.firewall' \
20+ --apply 'f: { inherit (f) allowedTCPPorts allowedUDPPorts extraCommands; }'
21+
22+## Stateful or not
23+
24+- **Hetzner Cloud Firewall:** stateful. Replies to outbound traffic are
25+ allowed automatically, so the inbound rules below plus the outbound
26+ catch-all are all it needs. None is active today.
27+- **netcup firewall policy:** behaves as stateless for UDP (observed
28+ 2026-10-08). Replies to the box's UDP queries arrive on a random high port
29+ and were dropped until rules allowed them by source port. TCP replies get
30+ through. So a netcup host also needs the [UDP reply rules](#udp-replies).
31+
32+## Per host
33+
34+"Opened by" is the module that opens the port in the NixOS firewall. Keep the
35+provider rule in step with it.
36+
37+### infra.mymood.at (netcup): moods, sovrn-metrics, sovrn-relay
38+
39+| Proto | Port | Source | Why | Opened by |
40+|---|---|---|---|---|
41+| TCP | 22 | any | SSH (deploys, admin) | servers `modules/base.nix` |
42+| TCP | 25 | any | Backup MX `mxb.eu.sovrn.at` (MX 20 for every sovrn domain) | sovrn `nix/modules/roles/relay.nix` |
43+| TCP | 80 | any | ACME HTTP-01 challenges, redirects to HTTPS | relay, metrics, moods proxy |
44+| TCP | 443 | any | moods, `infra.sovrn.at` (metrics, logs, alerting; cells push here) | sovrn `roles/metrics.nix`, moods `nix/modules/proxy.nix` |
45+| UDP | 443 | any | HTTP/3 for the same sites | moods `nix/modules/proxy.nix` |
46+| UDP | any | source port 53, 123, 24441 | Replies to the box's DNS, NTP and pyzor queries; see [UDP replies](#udp-replies) | none (needed only because netcup's policy doesn't track UDP) |
47+
48+netcup's policy as set on 2026-10-08 matches this table. Its default policy
49+also blocked outbound TCP 25, 465 and 587; that rule was deleted on
50+2026-10-06 (plan, "infra.mymood.at: network"). If netcup re-applies its
51+default, Alertmanager's mail (Resend, 587) stops.
52+
53+### mx99.eu.sovrn.at (Hetzner): sovrn-cell
54+
55+No provider firewall today. If one is added, every sovrn cell needs:
56+
57+| Proto | Port | Source | Why | Opened by |
58+|---|---|---|---|---|
59+| TCP | 22 | any | SSH | servers `modules/base.nix` |
60+| TCP | 25 | any | MX 10 for the cell's domains | sovrn `nix/modules/roles/cell.nix` |
61+| TCP | 80 | any | ACME HTTP-01, redirects | sovrn `roles/cell-edge.nix` |
62+| TCP | 443 | any | sovrnd, PDS origins (`*.pds1.eu.sovrn.at`), on-demand TLS | sovrn `roles/cell-edge.nix` |
63+| TCP | 465, 587 | any | Submission (implicit TLS, STARTTLS) for users' mail clients | sovrn `roles/cell.nix` |
64+| TCP | 143, 993 | any | IMAP (STARTTLS, implicit TLS) | sovrn `roles/cell.nix` |
65+| TCP | 2525 | the relays only (`152.53.243.113`, `2a0a:4cc0:2000:84c4:c4a5:5fff:fe80:c760`) | Relay ingress: the backup relay re-delivers queued mail here, unauthenticated, so it must never be open to anyone else | sovrn `roles/cell.nix` (`sovrn.cell.relaySources`, from the inventory) |
66+
67+ManageSieve (4190) and Stalwart's HTTP (8080, loopback) stay closed. The
68+relay addresses come from `hosts.json`; a generated rule should read them
69+from there as well, so a moved relay updates both layers.
70+
71+### infra.rtw.run (Hetzner): forge
72+
73+No provider firewall today.
74+
75+| Proto | Port | Source | Why | Opened by |
76+|---|---|---|---|---|
77+| TCP | 22 | any | SSH | servers `hosts/infra-rtw-run/services.nix` |
78+| TCP | 23231 | any | soft-serve: git over SSH | servers `roles/forge/default.nix` |
79+| TCP | 9418 | any | git daemon (public read-only clones) | servers `roles/forge/default.nix` |
80+| TCP | 80, 443 | any | soft-serve HTTP and kilimanjaro.io behind Caddy, ACME | servers `roles/forge/default.nix` |
81+
82+## UDP replies
83+
84+Needed only behind a firewall that doesn't track UDP (netcup). Each source
85+port and what breaks without it. None of these fails loudly.
86+
87+| Source port | Replies to | Without it |
88+|---|---|---|
89+| 53 | DNS: the host's systemd-resolved, and sovrn's unbound (Stalwart's resolver, recursing to authoritative servers anywhere) | unbound gets no answers, so Stalwart has no DNS and SPF, DMARC, DNSBLs and DANE fail. resolved still works through netcup's own resolvers, which hides the problem. |
90+| 123 | NTP (systemd-timesyncd) | The clock never syncs and drifts. The shared box was 31 s behind until this rule was added (2026-10-08). |
91+| 24441 | pyzor (Stalwart's spam filter queries `public.pyzor.org`) | Every inbound message waits out a 5 s timeout in the spam filter (bug d4988da). |
92+
93+A future service that queries something over UDP adds its server's port
94+here. Allowing inbound UDP to the local ephemeral range (32768-60999,
95+`net.ipv4.ip_local_port_range`) would cover everything at once, but it is
96+broader than needed while the provider can match on source port.
97+
98+## Outbound (reference)
99+
100+Covered by the catch-all. This list is for checking a provider's own
101+outbound blocks (netcup's default policy blocked SMTP ports; Hetzner blocks
102+some on new accounts) or for narrowing the catch-all one day:
103+
104+| Proto | Port | From | To | Why |
105+|---|---|---|---|---|
106+| UDP+TCP | 53 | all | any | DNS; unbound recurses to authoritative servers directly |
107+| UDP | 123 | all | NTP pool | Time |
108+| UDP | 24441 | cells, relay | `public.pyzor.org` | Spam filter |
109+| TCP | 443 | all | any | ACME (Let's Encrypt), Nix caches, R2 (Litestream, backups), Stalwart's spam-rule and ASN/geo downloads (GitHub), telemetry to `infra.sovrn.at`, SMTP2GO and other APIs |
110+| TCP | 2525 | cells | `mail-eu.smtp2go.com` | All outbound mail goes through SMTP2GO, which DKIM-signs and delivers (so DANE towards recipients is applied by SMTP2GO, not by the cells) |
111+| TCP | 2525 | relay | the cells | Re-delivery of queued mail to the cells' relay ingress |
112+| TCP | 587 | infra.mymood.at | `smtp.resend.com` | Alertmanager mail |
113+
114+Nothing needs outbound TCP 25: the cells send through SMTP2GO, and the relay
115+delivers only to the cells on 2525. Hetzner blocks outbound 25 and 465 on new
116+Cloud accounts until asked, so this layout keeps working there.
117+
118+## Checking
119+
120+From the host, after changing a provider rule:
121+
122+ # UDP replies (DNS, IPv4 and IPv6): expect status NOERROR, not "timed out"
123+ nix shell nixpkgs#dig -c dig +time=3 +tries=1 @198.41.0.4 . SOA
124+ nix shell nixpkgs#dig -c dig +time=3 +tries=1 @2001:503:ba3e::2:30 . SOA
125+ # NTP: NTPSynchronized=yes and a packet count above 0
126+ timedatectl show -p NTPSynchronized; timedatectl timesync-status
127+ # Stalwart's resolver (sovrn hosts): expect an answer within a second
128+ host -p 5335 example.com 127.0.0.1
129+
130+For pyzor, send a test message and compare the Stalwart log's DMARC and
131+`queue.message-queued` timestamps: about 0.2 s with a warm cache, 5 s more if
132+pyzor's replies are dropped. For inbound ports, test from outside (for
133+example check-host.net for TCP 25 and 443).