fleet-migration-plan.md
Fleet migration: sovrn, moods and servers → one Colmena fleet
Status: Migration complete (2026-10-08). Every host is deployed by this repository from the operator’s machine: infra.rtw.run (forge), infra.mymood.at (moods, sovrn’s metrics stack infra.sovrn.at and backup relay mxb.eu.sovrn.at), mx99.eu.sovrn.at (arm64 staging cell). sovrn’s Ansible is gone; recovery is automatic and rehearsed.
Goals
- infra.rtw.run on the standard method. Move it off agenix, the
Makefile and
nixos-rebuildonto Colmena and thesecretstool, and bring its packages up to date (it hasn’t been updated since June). It stays a separate box: it’s the base infrastructure for every project (personal git hosting now, a CI runner later). - Consolidate onto the netcup box.
infra.mymood.atruns moods today. Add sovrn’s metrics box and sovrn’s backup mail relay to it instead of building two more servers.
One Colmena hive, living in this repo, deploys every host. sovrn and moods keep developing independently (their own flakes, VM tests, devenv), and the fleet composes them onto hosts.
Principles
- One hive owns each host. Colmena deploys a whole system; two hives pointed at one box would each erase the other’s services. Once a host is in the fleet, only the fleet deploys it.
- Projects export, the fleet composes. sovrn and moods export overlays
and NixOS modules; they never import the fleet. The fleet pins each
project in
flake.lock, so “what is deployed” is a commit in this repo. - Nothing is updated implicitly. No
nix flake updateas a side effect of deploying. Bumping a project or nixpkgs is its own commit. - Prove no-ops before changing behaviour. Every adoption and refactor
first produces a system closure identical to what’s running (checked with
nvd diff), then a separate deploy makes the actual change. - One behavioural change per deploy, each with a rollback (previous
generation, or
colmena applyfrom the previous commit).
Phase 0 findings (2026-10-06)
Live hosts
Only two hosts are live. sovrn’s Ansible hosts (mx99.eu.sovrn.at, the
staging cell, and infra.sovrn.at, the metrics box) were torn down during
sovrn’s move to Colmena. They get rebuilt from scratch on the fleet, so there
is no data to migrate.
| Host | Provider | Size | Runs | OS |
|---|---|---|---|---|
infra.rtw.run |
Hetzner | 2 vCPU, 3.7 GiB, 38 GB | soft-serve 0.8.5, Caddy 2.10.0, pgit | NixOS 25.05 (ac62194, 2026-01-02), generation 68, up 114 days |
infra.mymood.at |
netcup | 4 vCPU, 7.75 GiB, no swap, 125 GB | moods (service + embedder), Caddy | NixOS 26.11pre (7a0f122), deployed by moods’ hive |
infra.rtw.run
- Running config is commit
8fd5074(2026-06-09,nixos-25.05locked toac62194c3917d5f474c1a844b6fd6da2db95077d). HEAD’sc3ca539(moves tonixos-26.05, locka0374025…) was never deployed. The Phase 2 baseline must use8fd5074’s lock to be a no-op. - No pending generation. Current, booted and default system are all generation 68, so an unplanned reboot changes nothing.
- stateVersion is 25.11 in the flake since 2026-04-14 (commit
89c8450), and every deployed generation since has used it. Keep 25.11. The leftover nixos-infect files in/etc/nixos(stateVersion 23.11, channelnixos-25.11) are unused; delete them in 2.7. - Root SSH keys. The hetzner key (
~/.ssh/hetzner, same assovrn/nix/keys/admins.pub) was meant to be added withssh-copy-id, but it never was: the “verified” logins used the scooter key (see 2.3). 2.2 step 5 put it in the config instead. sshd here reads both~/.ssh/authorized_keysand/etc/ssh/authorized_keys.d/%u(authorizedKeysInHomedirevaluates to true);/root/.sshis empty.
infra.mymood.at: headroom
Measured:
| RAM | CPU | Disk | |
|---|---|---|---|
| moods-embedder | 1.75 GB, flat (peak = current) | ~1.4 cores average (447k CPU-s over 3.8 days) | model in the Nix store |
| moods (BEAM) | 1.1 GB cgroup, includes SQLite page cache | ~0.03 cores | /var/lib/moods 3.6 GB |
| Caddy + system | ~0.2 GB | – | – |
| Box | 2.35 GB used, 5.6 GB available | load ~1.1, CPU pressure ~1.4% | 5.9 / 125 GB |
Estimated additions (nothing to measure yet):
| RAM | CPU | Disk | |
|---|---|---|---|
| sovrn metrics (VictoriaMetrics, VictoriaLogs, 2× vmalert, Alertmanager, healthz), 1–3 cells | 0.5–1 GB | <0.2 core | VictoriaLogs capped at 25 GiB; VictoriaMetrics 90d, a few GB |
| sovrn relay (Stalwart store-and-forward + route-sync) | 0.2–0.4 GB | ~0 idle | queue, small |
| Projected total | ~3.5–4 of 7.75 GB | ~1.7 of 4 cores | ~35–40 of 125 GB worst case |
It fits, with two guard rails (Phase 3.3): swap (there is none, so a spike
goes straight to the OOM killer) and limits on moods-embedder
(MemoryHigh/MemoryMax, lower CPUWeight) so an embedding backlog can’t
starve Stalwart or Alertmanager.
infra.mymood.at: is co-locating safe?
- moods + metrics: yes. No port overlap, same nixpkgs rev, both only serve HTTP behind Caddy.
- + relay: yes, given the network prerequisites below.
- The relay covers Hetzner cells from a netcup box: a backup MX at a different provider from what it backs up is better than one beside them.
- Ports: Stalwart takes 25 (public) and 8080⁄18080 (loopback); nothing else on the box uses them.
- Its certificate comes from ACME HTTP-01 through a Caddy
http://vhost, the same port-80 challenge proxy cells use. That merges with moods’ and metrics’ vhosts. - It runs sandboxed as its own user with secrets in
/var/lib/sovrn-keys; moods runs asmoods, hardened, and can’t read them. - Blast radius: if the box dies, moods, alerting and the backup MX go
together. Losing the backup MX only matters if a cell fails at the same
time, and the external Larm monitor on
https://infra.sovrn.at/healthzstill notices the box is down.
- Cells stay on their own hosts. A cell runs mail on 25/465/587/143/993, a
catch-all
:443with on-demand TLS for tenant PDS hosts, and a per-host Stalwart plan. The fleet enforces this with an assertion.
infra.mymood.at: network
- Outbound SMTP: fixed. netcup’s default firewall policy blocked 25, 465
and 587 outbound. That rule was deleted on 2026-10-06; after that, port 25
connected over IPv4 (Gmail, Fastmail, iCloud) and IPv6 (Fastmail), and
Gmail answered
220 mx.google.com ESMTP. Resend 465⁄587 are open, so Alertmanager can keepsmtp.resend.com:587. If netcup ever re-applies its default policy, the relay silently stops delivering: the relay’s health check should include an outbound-25 probe (Phase 5). - Inbound 25: not yet tested (needs a listener). Check on the relay’s first deploy.
- Reverse DNS: IPv4 PTR is
infra.mymood.at; IPv6 has none. Set both to the relay’s mail name (mxb.eu.sovrn.at) in netcup’s SCP before the relay goes live. moods and metrics only serve HTTP and don’t care.
Where things are today
| sovrn | moods | servers | |
|---|---|---|---|
| Deploy | Colmena hive (Justfile.nix), no hosts; Ansible in deployment/ (hosts torn down) |
Colmena hive (Justfile) |
Makefile → scripts/update.sh → nixos-rebuild on the target |
| NixOS hosts | none (nix/hosts.json is {}) |
infra.mymood.at |
infra.rtw.run |
| nixpkgs | unstable 7a0f122 (Stalwart 0.16) |
unstable 7a0f122 (kept equal to sovrn) |
running 25.05 ac62194; repo HEAD says 26.05 |
| Secrets | secrets CLI → Colmena keys in /var/lib/sovrn-keys |
same, /var/lib/moods-keys |
agenix, encrypted to the host SSH key |
| Install | nixos-anywhere + disko, host key from the store | same, plus netcup | cloud-init + nixos-infect, hardware.nix fetched afterwards |
| Caddy | services.caddy |
services.caddy |
hand-written systemd.services.caddy + Caddyfile |
| Relay | nix/modules/roles/relay.nix is a placeholder; Ansible role deployment/roles/stalwart_backup (install, bootstrap, routesync) to port from |
– | – |
Target
| Host | Roles | Notes |
|---|---|---|
infra.rtw.run (Hetzner) |
forge (later ci) |
Separate base infrastructure. In place on its legacy layout first; optional rebuild later (2.8) |
infra.mymood.at (netcup) |
moods, sovrn-metrics, sovrn-relay |
The shared box. DNS names infra.sovrn.at and mxb.eu.sovrn.at point here |
| sovrn cells (Hetzner, new) | sovrn-cell |
Exclusive. mx99 rebuilt first as staging |
Cost: metrics and the relay each needed their own box in the Ansible design. Here they cost nothing extra; the fleet is two always-on boxes plus however many cells sovrn runs.
Shared box port map
| Owner | Ports |
|---|---|
| moods | 7071 (embedder), 7072 (service), loopback |
| sovrn metrics | 8428, 9428, 8880, 8881, 9093, 8090, loopback |
| sovrn relay | 25 public; 8080 (Stalwart http), 18080 (recovery mode), loopback |
| Caddy (shared) | 80, 443/tcp, 443/udp |
Caddy vhosts on the shared box: mymood.at, feeds.mymood.at (moods),
infra.sovrn.at (metrics), http://mxb.eu.sovrn.at (ACME challenge proxy
to Stalwart only).
Repository layout (this repo)
servers/
flake.nix inputs: nixpkgs (= sovrn/moods rev), nixpkgs-stable,
colmena v0.5.0, disko, pgit, sovrn, moods
flake.lock the deployed versions of everything
Justfile eval, build, diff-host, deploy, deploy-project,
check-secrets, known-hosts, sync-nixpkgs
(new-host, gen-host-secrets in Phase 4)
hosts.json inventory (below)
hive.nix inventory → Colmena hive; also nixosConfigurations;
fails evaluation if sovrn's or moods' nixpkgs differs
keys/admins.pub root SSH keys for standard hosts (same as sovrn/moods)
modules/
base.nix every standard host: SSH, firewall, resolved, nix GC, admin keys
vm.nix disko + GRUB + networkd (from moods)
hetzner.nix from moods
netcup.nix from moods
secrets.nix `servers.secrets`, /var/lib/servers-keys (moods' module, renamed)
fleet.nix `fleet.roles`; assertion: an exclusive role (sovrn-cell) is alone on its host
caddy.nix fleet-wide Caddy settings (ACME email); vhost-collision assertion still to come
roles/
forge.nix soft-serve, pgit hook, backup timer, Caddy vhosts
hosts/
infra-rtw-run/ legacy layout: hardware.nix, static networking
docs/
Inventory (hosts.json)
moods’ schema plus roles, nixpkgs and layout:
{
"infra.mymood.at": {
"provider": "netcup", "system": "x86_64-linux", "disk": "/dev/vda",
"ipv4": "152.53.243.113", "ipv4Prefix": 22, "ipv4Gateway": "152.53.240.1",
"ipv6": "2a0a:4cc0:2000:84c4:c4a5:5fff:fe80:c760/64",
"sshHostKey": "ssh-ed25519 …",
"roles": ["moods", "sovrn-metrics", "sovrn-relay"],
"settings": { "moods": {}, "sovrn": {} }
},
"infra.rtw.run": {
"layout": "legacy", "nixpkgs": "stable", "system": "x86_64-linux",
"ipv4": "178.104.201.207", "ipv6": "2a01:4f8:1c18:235d::1/64",
"sshHostKey": "ssh-ed25519 … (from identities/host/infra-rtw-run.pub)",
"roles": ["forge"]
}
}
rolespicks the modules:moods→moods.nixosModules.app,forge→./roles/forge.nix,sovrn-metrics→sovrn.nixosModules.metrics,sovrn-relay→sovrn.nixosModules.relay,sovrn-cell→sovrn.nixosModules.cell. Roles are also Colmena tags:colmena apply --on @moodsdeploys every host running moods.layout: "legacy"importshosts/<name>/(hardware.nix + hand networking) instead ofvm.nix+ a provider module. Only infra.rtw.run uses it.nixpkgs: "stable"selectsnixpkgs-stablethroughmeta.nodeNixpkgs(both existing hives already setnodeNixpkgs). Default: the shared rev.settings.<project>feeds per-host overrides into project options.
What sovrn and moods export
| Output | sovrn | moods |
|---|---|---|
overlays.default |
exists | exists |
nixosModules.services |
stalwart, sovrnd, zds, secrets, sovrn user (all off by default) | moods.nix, secrets.nix |
nixosModules.<role> |
metrics, relay, cell |
app (app.nix + proxy.nix) |
checks |
VM tests (stay; add relay) |
VM test (stays) |
Removed from the projects at the end (Phase 11): hosts.nix, hosts.json,
base.nix, provider modules, new-host/known-hosts/deploy recipes.
Secrets
One secrets store (SECRETS_DIR, ~/projects/data), one prefix per
project, one key directory per project on the host:
| Namespace | Store paths | On host |
|---|---|---|
sovrn.secrets |
sovrn/shared/…, sovrn/hosts/<fqdn>/… |
/var/lib/sovrn-keys |
moods.secrets |
moods/shared/…, moods/hosts/<fqdn>/… |
/var/lib/moods-keys |
servers.secrets |
servers/shared/…, servers/hosts/<fqdn>/… |
/var/lib/servers-keys |
Colmena’s deployment.keys is one attrset per host, and each key gets a
<attr>-key.service unit. On the shared box, sovrn’s metrics and relay and
moods all declare keys, so names must not collide: each project’s secrets
module prefixes the attribute (sovrn-metrics-password) and keeps the
on-host file name through Colmena’s name option. Confirmed in Colmena
0.5.0’s source: the <attr>-key path and service units are named after the
attribute, and the file after name. Services order on the unit name the
module exposes (.unit), never a hard-coded string. The fleet’s
servers.secrets already works this way (servers-<name>).
Within sovrn, metrics and relay share one key directory. Today their names
don’t overlap (relay: stalwart-recovery-password, backup-domains-token,
metrics-password; metrics: metrics-password-hash, pds-resend-api-key),
but a name declared by two sovrn roles on one host must have identical
settings, or the module system rejects it.
Toolkit used throughout
# What a host is running
ssh root@HOST readlink -f /run/current-system
# What the fleet would deploy
nix build ".#nixosConfigurations.\"HOST\".config.system.build.toplevel" -o result-HOST
# Compare (fetch the running closure first)
nix copy --from ssh-ng://root@HOST /nix/store/…-nixos-system-…
nix run nixpkgs#nvd -- diff /nix/store/…-nixos-system-… ./result-HOST
# Which units a deploy would touch, without touching them
colmena apply dry-activate --on HOST
# Deploy / deploy and reboot into it (build + boot + reboot)
colmena apply --on HOST
colmena apply --on HOST --reboot
# Roll back on the host
ssh root@HOST nixos-rebuild switch --rollback # or pick the generation in GRUB
Wrap these as just diff-host HOST and just deploy HOST in the fleet’s
Justfile. just deploy runs check-secrets HOST first, like sovrn and moods.
Phase 0: Survey and safety nets
Done:
- [x] Live hosts confirmed (only infra.rtw.run and infra.mymood.at)
- [x] infra.rtw.run’s running commit, nixpkgs rev, stateVersion, pending generations
- [x] infra.mymood.at headroom measured
- [x] netcup outbound 25/465/587 unblocked and verified
- [x] hetzner key on infra.rtw.run, verified with
ssh -i ~/.ssh/hetzner -o IdentitiesOnly=yes [email protected] true
Before each host’s first fleet deploy:
- [ ] Hetzner snapshot of infra.rtw.run; netcup snapshot of infra.mymood.at
- [ ]
ssh [email protected] systemctl start backupand confirmr2:soft-serve/is current - [ ] Decide CI remotes: the fleet references sovrn and moods as
git+file:///home/btburke/projects/<p>from the workstation; CI needs remote URLs
Phase 1: Fleet skeleton (no deploys) — done 2026-10-06
Built as below, with these differences from the original sketch:
- sovrn and moods are
git+file:///…?ref=HEAD. Plaingit+filereads the working tree, which Nix refuses to lock while a repo has uncommitted changes (sovrn did).ref=HEADis the last committed change (jj:@-). Locked: sovrnf63de9f, moods476bee5. - The nixpkgs check lives in
hive.nix(evaluation fails with instructions), the exclusive-role check inmodules/fleet.nix.caddy.nixwaits for Phase 8, when two projects first share Caddy. nixpkgs.flake.sourceis set only for legacy hosts:nixosSystemset it for infra.rtw.run, moods’ hive never did for infra.mymood.at.- The old entry point
nixosConfigurations.infra-rtw-runis kept (onnixpkgs-stable, i.e. what’s running) until the Makefile goes in 2.3. Don’t use the Makefile anyway:update.shbuilds on the server fromgit+https://git.kilimanjaro.io/servers, where thegit+fileinputs can’t resolve.
Verified (evaluation only, nothing built or deployed):
- [x]
nixosConfigurations.infra-rtw-runevaluates to/nix/store/r46cnzkj…-nixos-system-infra-25.05.20260102.ac62194, the system infra.rtw.run runs - [x] with infra.rtw.run temporarily in
hosts.json(legacy), the Colmena node evaluates to the same path: step 2.2 is a no-op - [x] with infra.mymood.at temporarily in
hosts.jsonand moods’ current modules wired in as a role, the node evaluates to/nix/store/jh2lndcc…-nixos-system-infra-26.11.20260928.7a0f122, what it runs and what moods’ hive builds: the copied base/vm/netcup modules match, so step 3.2 can be a no-op too - [x] overriding the fleet’s nixpkgs to another rev fails evaluation with the “bump them together” message
- [x]
fleet.roles = [ "sovrn-cell" "moods" ]fails its assertion - [x] in
nix develop: colmena 0.5.0, nvd 0.2.4;just evalandjust check-secretspass on the empty inventory
The inventory is back to {}; hosts are added in 2.2 and 3.2.
Original sketch:
- Flake inputs:
nixpkgs=github:NixOS/nixpkgs/7a0f122f5090cf4c2ade2a13a0e229d4e19ba71f(sovrn’s and moods’ rev)nixpkgs-stable=github:NixOS/nixpkgs/ac62194c3917d5f474c1a844b6fd6da2db95077dat first: exactly what infra.rtw.run runs, so 2.2 is a no-op. It moves tonixos-26.05in 2.5.colmenav0.5.0,disko,sovrn,moodspgitandagenixlocked to the revs in8fd5074’sflake.lock(agenixis removed in 2.4)
devShells.default: colmena, nixos-anywhere, nvd, jq, age, just. (secretsis~/bin/secrets; the shell just needs it on PATH.)hive.nix: inventory → hive, modelled on moods’nix/hosts.nix(nodeNixpkgs,nodeSpecialArgs,versionSuffix/revision,targetHost= name,targetUser = "root"). ExposecolmenaHiveandnixosConfigurations = colmenaHive.nodes.modules/: copybase.nix,vm.nix,hetzner.nix,netcup.nix,secrets.nixfrom moods;base.nixauthorizesadmins.pubfor root. Not used by infra.rtw.run (legacy layout).nixpkgs-check.nix: assertinputs.sovrn.inputs.nixpkgs.rev == inputs.nixpkgs.revand the same for moods, with a message that says to bump them together.just eval(Colmena eval of every node) passes with an empty inventory.
Pinning nixpkgs to the project rev keeps sovrn’s and moods’ promise that dev builds are host builds: the fleet applies the projects’ overlays to the same nixpkgs, so the store paths match.
Phase 2: infra.rtw.run → Colmena, secrets tool, current packages
In place, no reinstall. Each numbered deploy is separate. Order: adopt unchanged → keys → agenix out → upgrade → Caddy. The upgrade comes before the Caddy conversion so that conversion is written against the Caddy module version you’ll keep.
2.1 Put the secrets in the store
Done 2026-10-06: all four entries decrypt; pushover.env has both
PUSHOVER_ variables, rclone.conf the [r2] remote, the hash is a crypt
string, and the host key’s public half matches identities/host/infra-rtw-run.pub.
The plaintext sources (~/data) aren’t on the workstation, so read each
value off the live host and pipe it straight into the store; nothing lands on
local disk.
ssh [email protected] cat /etc/pushover.env \
| secrets encrypt servers/shared/pushover.env
ssh [email protected] cat /root/.config/rclone/rclone.conf \
| secrets encrypt servers/hosts/infra.rtw.run/rclone.conf
ssh [email protected] cat /etc/btburke-password \
| secrets encrypt servers/hosts/infra.rtw.run/btburke-password-hash
ssh [email protected] cat /etc/ssh/ssh_host_ed25519_key \
| secrets encrypt servers/hosts/infra.rtw.run/ssh_host_ed25519_key
If the R2 credentials in rclone.conf are the same as the store’s top-level
rclone.conf.age, replace the host entry with a symlink to it instead
(secrets follows links).
Check: secrets decrypt servers/shared/pushover.env | grep -c PUSHOVER_ → 2.
2.2 Adopt the host unchanged
- Add
infra.rtw.runtohosts.json(layout: legacy,nixpkgs: stable,roles: []for now,sshHostKeyfromidentities/host/infra-rtw-run.pub). - Its node imports exactly what
8fd5074’sflake.niximports:agenix.nixosModules.defaultand./hosts/infra-rtw-run, with the same specialArgs (flakeRoot,flakeHostName = "infra-rtw-run",pgit). Setsystem.nixos.versionSuffix/revisionlike moods’hosts.nix, so the system label matches anixosSystembuild. deployment:targetHost = "infra.rtw.run",targetUser = "root",buildOnTarget = false. Builds moved to the server in6b79956; only setbuildOnTarget = trueif the reason for that still applies.just known-hostspins the key for both name and IP.
Done 2026-10-06: steps 1–4; just diff-host infra.rtw.run built the system
locally and printed “identical” (r46cnzkj…); just known-hosts pinned the
name (already that key) and the IP.
Verify it’s a no-op: just diff-host infra.rtw.run prints “identical”
(already confirmed at evaluation in Phase 1). Then
colmena apply dry-activate --on infra.rtw.run lists no units to restart.
- Root keys, separate commit. Add the hetzner key next to the scooter key
in
hosts/common/users.nix(orkeyFiles = [ admins.pub ]), so it lands in/etc/ssh/authorized_keys.d/root. Kept out of the adoption commit so the no-op check above stays clean.
2.3 Deploy 1: switch deploy tooling
Done 2026-10-06:
- [x] deploy of
olkkkkux(adoption): same systemr46cnzkj…, still generation 68, soft-serve and Caddy not restarted - [x] deploy of
uwpmokuu(root keys): generation 69;/etc/ssh/authorized_keys.d/roothas the scooter and hetzner keys; hetzner-key-only login works - [x] Makefile deleted; README/AGENTS point at the Justfile until 2.7
- [x]
remove the: there never was one.ssh-copy-idline~/.ssh/confighasHost infra.rtw.run→IdentityFile ~/.ssh/rtw(the scooter key), and ssh offers that file even with-i ~/.ssh/hetzner -o IdentitiesOnly=yes. Sossh-copy-id’s “is it already installed?” probe logged in with the scooter key and skipped the copy, and every “hetzner key only” check before deploy 2 also used the scooter key. sshd’s log shows the hetzner key’s first login at 06:02:23, right after deploy 2. A real test ignores the ssh config:ssh -F /dev/null -o UserKnownHostsFile=~/.ssh/known_hosts -i ~/.ssh/hetzner -o IdentitiesOnly=yes [email protected] true(passes since deploy 2).
Found during 2.3: leaked D-Bus daemons. Both activations printed
reloading user units for root... Failed to open dbus connection, and
[email protected] is failed with “Too many open files”. The post-receive
hook’s pgit/pgit-index runs autolaunch a session dbus-daemon that never
exits: 462 of them since 2026-06-23, each holding an inotify instance, which
exhausted root’s fs.inotify.max_user_instances (128). Must be fixed before
2.4, whose Colmena key units run inotifywait as root. See 2.3a.
colmena apply --on infra.rtw.run
Nothing on the box changes. From now on the Makefile must not be used:
delete the update, bootstrap, recover and reencrypt targets in the
same commit (or the whole Makefile; the rest goes in 2.7).
Then deploy the root-key commit from 2.2 step 5 and check:
- [x]
grep -c ICLloca /etc/ssh/authorized_keys.d/root→ 1 - [x] the hetzner key logs in on its own, tested with
ssh -F /dev/null(see below) - [ ] optional: to drop the scooter key from the config, first point
~/.ssh/config’sHost infra.rtw.runat~/.ssh/hetzner(Colmena connects through that block and uses~/.ssh/rtwtoday), deploy once with it, then remove the key
2.3a Stop the D-Bus leak (before 2.4)
Confirmed on the host: pgit-index with DBUS_SESSION_BUS_ADDRESS unset
leaves one more orphaned dbus-daemon; with it set to disabled: it leaves
none.
- Commit: the post-receive hook exports
DBUS_SESSION_BUS_ADDRESS=disabled:. Only the hook (under/etc/soft-serve/hooks) and/etcchange; nothing restarts. - Deploy:
just deploy infra.rtw.run. - Clean up the orphans (root session buses only; the system bus runs as
messagebuswith--system):pkill -u 0 -f -- 'dbus-daemon --syslog --fork --print-pid 4 --print-address 6 --session', thensystemctl reset-failed [email protected]. - Check (done 2026-10-06, generation 70):
- [x] orphans gone: 463 → 0 (the only root
dbus-daemonleft is root’s real user bus,--address=systemd:under[email protected]); the system bus (messagebus) untouched - [x]
[email protected]active,busctl --userworks, no failed units; inotify instances on the box: 13 in total - [x] a repeat deploy activates without
Failed to open dbus connection - [x] a push (the servers repo, 06:10) rebuilt the site index
(13 repositories) and left no root
dbus-daemon
- [x] orphans gone: 463 → 0 (the only root
Note for later: pkill -f <pattern> over ssh … '<command>' also matches
the remote shell running the command (its argv contains the pattern) and
kills it mid-way. Use pgrep/kill on PIDs, or a pattern trick like
[d]bus-daemon.
2.4 Deploy 2: agenix → Colmena keys
Implemented 2026-10-06 (commit “agenix → Colmena keys”): hosts/common/secrets.nix
imports modules/secrets.nix and declares pushover.env;
hosts/infra-rtw-run/secrets.nix declares rclone.conf and
btburke-password-hash; services order after <key>.unit. The agenix input,
the legacy nixosConfigurations.infra-rtw-run output and agenix in the dev
shell are gone (secrets/*.age stay in git until 2.7). Root also gets
RCLONE_CONFIG in its environment, so interactive rclone and recover.sh
find the config.
In hosts/infra-rtw-run/:
servers.secrets = {
"pushover.env".scope = "shared";
"rclone.conf".scope = "host";
"btburke-password-hash".scope = "host";
};
All three are root-owned, so Colmena uploads them before activation; the password hash has to exist when the users step of activation reads it.
| Was | Becomes |
|---|---|
age.secrets.*, agenix module, age.identityPaths |
removed |
hashedPasswordFile = config.age.secrets.btburkePassword.path |
config.servers.secrets."btburke-password-hash".path |
EnvironmentFile = "/etc/pushover.env" (soft-serve, backup, system-monitor) |
the pushover.env key path; wants/after its key unit |
rclone reads /root/.config/rclone/rclone.conf |
backup.service: environment.RCLONE_CONFIG = the rclone.conf key path |
recover.sh uses the default rclone config |
export RCLONE_CONFIG in it too, and fix its failure branch, which calls an undefined log |
Users: servers doesn’t set users.mutableUsers = false, so NixOS keeps the
existing password of an existing user, and swapping the hash file’s source
doesn’t change btburke’s password. Leave mutableUsers alone on this host.
Verify before deploying: just diff-host shows agenix leaving and only the
soft-serve, backup, system-monitor and user units changing.
just check-secrets infra.rtw.run passes.
colmena apply --on infra.rtw.run --reboot
After it’s back:
Deployed 2026-10-06 with --reboot (generation 71, booted). A backup ran
just before, on the agenix setup.
- [ ]
ssh [email protected], thensudo trueasks for and accepts the password (needs the password: by hand) - [x]
/var/lib/servers-keys(0711) holds three root-owned 0400 files; the threeservers-*-keyunits are active; no failed units - [x]
systemctl start backupsucceeds; Pushover answered"status":1 - [x]
systemctl start system-monitorruns without the missing-token warning - [x] soft-serve read
/var/lib/servers-keys/pushover.envand sent its start notification; clone over HTTPS andgit://, andgit ls-remoteover SSH (23231) work; https://kilimanjaro.io → 200 - [x] root’s login shell has
RCLONE_CONFIG=/var/lib/servers-keys/rclone.conf - [x] dead agenix symlinks removed (
/etc/pushover.env,/etc/btburke-password,/root/.config/rclone/rclone.conf)
Rollback: the previous generation in GRUB, or colmena apply from the
commit before. agenix’s ciphertext is still in git until 2.7.
2.5 Deploy 3: upgrade 25.05 → 26.05
soft-serve dry run (done 2026-10-06). 26.05 takes soft-serve from 0.8.5 to
0.11.6 (golang.org/x/crypto v0.49.0). The production data (/var/soft/data,
66 MB) was copied off the box and served locally by the 0.11.6 build from
nixos-26.05, with the box’s config.yaml, local-only ports and the global
hooks replaced by a test hook:
- [x] key exchange: 0.8.5 offers nothing post-quantum (best
curve25519-sha256, so OpenSSH 10 prints the “not using a post-quantum key exchange” warning on every push); 0.11.6 negotiatesmlkem768x25519-sha256, no warning - [x] database: no new migrations (stays at 3: create tables, webhooks,
migrate_lfs_objects);
PRAGMA integrity_checkok before and after - [x] 14 repos, 1 admin user,
anon-access=read-only,allow-keyless=true,moodsstill private; anonymous HTTP to it refused - [x] same SSH host key (
SHA256:mxf8tdRp…), so clients won’t warn - [x] clone over SSH, HTTP and
git://; push to an existing repo and push creating a new one; untouched repos’ refs identical afterwards - [x] the global
hooks/post-receivestill runs on push (the pgit rebuild depends on it); per-repo hooks call$SOFT_SERVE_BIN_PATH, so they don’t pin the old store path - [x]
config.yaml: every key 0.11.6’s default config has is present; the only extra one,initial_admin_keys, is still accepted (and only used when a new database is created) - [x] no errors or warnings in the server log
The copy (which included soft-serve’s private host key and the private repo) was deleted afterwards.
Rollback is still GRUB plus the step-2 backup, but since the database schema doesn’t change, 0.8.5 can read what 0.11.6 leaves behind.
This is goal 1’s package update: soft-serve, Caddy, the kernel and everything else move to 26.05. It skips 25.11.
- Read the 25.11 and 26.05 release notes for what this host uses through
NixOS modules:
openssh,users,networking(scriptednetworking.interfaces+ dhcpcd),security.sudo,programs.neovim, GRUB. soft-serve and Caddy are hand-written units here, so their module changes don’t apply, but check soft-serve’s own changelog from 0.8.5 for config or database migrations. - Fresh backup:
systemctl start backup. - Point
nixpkgs-stableatgithub:NixOS/nixpkgs/nixos-26.05,nix flake update nixpkgs-stable. Also bumppgitif wanted (it’s built with the host’s Go, so it changes anyway). just diff-host infra.rtw.run: review the version changes (soft-serve, Caddy, kernel, systemd, openssh). Fix eval warnings and errors from renamed options.-
colmena apply --on infra.rtw.run --reboot - Check (deployed 2026-10-06 with
--reboot, generation 72):- [x]
nixos-version26.05.20261006.b253099; soft-serve 0.11.6, Caddy 2.11.4, OpenSSH 10.5p1, kernel 6.18.55, systemd 260.4 - [x] SSH as root (hetzner key and scooter key)
- [x] IPv4: default route via 172.31.1.1, egress works. IPv6: see below
- [x] git over SSH negotiates
mlkem768x25519-sha256(no post-quantum warning);git://and HTTPS clone work - [x] https://kilimanjaro.io 200, www → 301; https://git.kilimanjaro.io/
is 404 straight from soft-serve (it has no page at
/; repo paths return 200), HSTS header present - [x] a push returns at once (after 2.5b) and the background build regenerates the site (index rebuilt 07:12:41, 13 repos); no D-Bus orphans
- [x]
systemctl --failedis empty; backup and system-monitor run; no rootdbus-daemonorphans
- [x]
Caught before deploying: the default route. 26.05’s scripted networking
only installs networking.defaultGateway on an interface it’s named for, or
whose subnet contains the gateway. 172.31.1.1 is outside the /32, and
defaultGateway had no interface, so the 26.05 build had no IPv4 default
route at all; booted, the box would have been unreachable except through
Hetzner’s console. (25.05’s network-setup.service added it unconditionally;
26.05 removed that unit.) Fixed by defaultGateway = { address = "172.31.1.1";
interface = "enp1s0"; }, checked in the generated network-addresses-enp1s0
script before deploying. The release note says “implementation details”;
this is the kind of change only the generated units show.
Found: IPv6 has never worked on this box. There is no IPv6 default route
(defaultGateway6 is unset; only fe80::1/128 is routed), so outbound IPv6
fails and the AAAA record for infra.rtw.run points at an address that can’t
answer anyone off-link. kilimanjaro.io’s AAAA records are Cloudflare’s and
git.kilimanjaro.io has none, so the services aren’t affected. Fix
separately (2.5a).
2.5b Slow pushes after the upgrade (fixed 2026-10-06)
After 2.5 every push waited for the whole pgit site build (about a minute)
even though the hook backgrounds it. Reproduced locally with a stand-in
15-second job: 0.2 s per push on soft-serve 0.8.5, 16.2 s on 0.11.6. The
background job inherited fd 5, a pipe whose other end is held by the
soft serve server process, which only ends the push at EOF on it. The fix
in the hook: the background job closes every descriptor above 2 (and takes
stdin from /dev/null) before doing anything. Locally: 0.2 s on both
versions, every background build still completes.
2.5a IPv6 default route (done 2026-10-06, generation 74)
networking.defaultGateway6 = { address = "fe80::1"; interface = "enp1s0"; };
(Hetzner’s IPv6 gateway).
Not deployed with switch: changing network-addresses-enp1s0 restarts it,
and its stop step deletes every address and route on the interface (IPv4
included) before the start step adds them back. Instead:
- Added the route by hand on the running box (
ip -6 route replace default via fe80::1 dev enp1s0): IPv6 egress worked, and from infra.mymood.at (which has IPv6) inbound SSH, HTTP (308 to HTTPS) and git SSH on 23231 answered over IPv6. - Deployed with
--reboot.
After the reboot: both default routes present, curl -4 and curl -6 to
example.com → 200, inbound IPv6 HTTP from infra.mymood.at → 308, no failed
units. (Go’s SSH server on 23231 waits for the client’s version string
before sending its own, so a bare banner probe looks like a dead port.)
Rollback: generation 69⁄70 in GRUB. soft-serve migrates its database forward on start, so rolling back across a soft-serve upgrade may need the R2 backup from step 2.
2.6 Deploy 4: Caddy onto services.caddy
Findings before implementing (2026-10-06):
- The old unit ran as root with
HOME=/root, so Caddy’s storage (three certificates and the ACME account) is in/root/.local/share/caddy;/var/lib/caddywas empty. Renewals worked: the origin certificates were renewed in August and are valid to about 2026-11-10. - kilimanjaro.io and www are proxied by Cloudflare (visitors see Cloudflare’s certificate); only git.kilimanjaro.io points at the box.
- The NixOS module runs Caddy as
caddy(home/var/lib/caddy, so storage/var/lib/caddy/.local/share/caddy), logs every site to/var/log/caddy/access-<host>.logand sets the default logger to ERROR.
Implemented in roles/forge.nix (infra.rtw.run gets roles: ["forge"]):
the three sites verbatim; logFormat = null per site and level INFO
globally to keep today’s logging (no access logs, renewals in the journal);
a one-time caddy-storage-migrate unit copies the old storage into the new
location before Caddy starts, so nothing is reissued. caddy adapt of the
old and new Caddyfiles gives identical JSON apart from the source path and
the explicit INFO level (the old default).
- Find where the current certificates live. The hand-written unit runs as
root with
ProtectSystem = "strict", writable/var/lib/caddyand noXDG_DATA_HOME:ssh [email protected] 'ls -R /var/lib/caddy | head; ls /root/.local/share/caddy'. - Write
roles/forge.nixwithservices.caddy.virtualHostsforgit.kilimanjaro.io,www.kilimanjaro.ioandkilimanjaro.io, eachextraConfigtaken from the current Caddyfile verbatim. Drop therequires = soft-serve.servicecoupling so the static site doesn’t go down with soft-serve. - Remove
systemd.services.caddy,environment.etc."caddy/Caddyfile",caddyfrom systemPackages, and the/var/lib/caddytmpfiles rule. - Certificates: either copy the existing ones into the module’s data
directory (owned by
caddy), or let Caddy issue fresh ones; three names are far under Let’s Encrypt’s limits, at the cost of a few seconds of TLS errors on the first requests.
colmena apply --on infra.rtw.run
Deployed 2026-10-06 (plain switch, generation 75):
- [x]
caddy-storage-migrateran once;/var/lib/caddy/.local/share/caddyholds the three certificates, owned bycaddy; Caddy runs ascaddy - [x] certificates reused, not reissued: the origin serials for all three names are unchanged
- [x] https://kilimanjaro.io 200, www → 301 kilimanjaro.io, git.kilimanjaro.io/ 404 (soft-serve; repo paths 200); HTTPS clone works
- [x] HSTS, nosniff, X-Frame-Options present; no
Serverheader; gzip; CSS gets the immutable cache header - [x] Caddy’s journal: normal startup, no ACME activity or errors
Seen during activation, unrelated: efi.mount started. The disk carries an
EFI system partition (sda15) from Hetzner’s cloud image, and systemd 260’s
GPT auto-generator automounts it at /efi on first access. Harmless.
Left for 2.7: delete /root/.local/share/caddy and /root/.config/caddy,
then drop the caddy-storage-migrate unit.
2.7 Move the rest into the forge role; remove the old tooling
- Move soft-serve, its config, the post-receive hook, the backup timer and
the start notification from
hosts/infra-rtw-run/services.nixintoroles/forge.nix; setroles: ["forge"].nvd diffshould show no change. - Fix the typo in
config.yaml:ssh.public_urlisssh://git.kilimarjaro.io:23231(should bekilimanjaro). - Pull the site-generation loop out of the post-receive hook into a
forge-rebuild-sitecommand and have the hook call it./var/www/codeisn’t backed up, so after any restore this is how it comes back. - Delete
Makefile,scripts/,state/,secrets/,identities/,docs/agenix-rekey*.md, the agenix input and the devShell’s agenix/bash setup. Keepzone.dbandrecover.sh(move underroles/forge/). RewriteREADME.mdandAGENTS.mdfor the Colmena workflow. - On the host:
rm -r /etc/nixos(nixos-infect leftovers) andnix-channel --remove nixos.
Done 2026-10-06 (generation 76):
- Move: soft-serve, hook, backup and Caddy into
roles/forge/(withconfig.yaml,post-receive,backup.sh,recover.sh); built to the identical system before any change. forge-rebuild-site(roles/forge/rebuild-site.sh, awriteShellApplicationwith pgit and sqlite) holds the site build; the hook only backgrounds it. Ran by hand on the box: 13 repos + index in 52 s, no D-Bus leak.recover.shnow runs it after a restore.config.yaml:kilimarjaro→kilimanjaro; soft-serve restarts when the file changes.- Removed
caddy-storage-migrateandnetwork-autodetect.nix(its only reader wasscripts/bootstrap-host.sh);jqstays on PATH viahosts/common/base.nix. - On the box: deleted
/etc/nixos(nixos-infect), root’snixos-25.11channel,/root/.local/share/caddyand/root/.config/caddy(identical to the migrated copies) and/root/result(pinned the 25.05 system). - Repo: deleted
scripts/,secrets/(agenix ciphertext),identities/,docs/agenix-rekey*.md,docs/TESTING.md;.gitignoretrimmed; README, AGENTS,hosts/common/README.md,hosts/infra-rtw-run/README.mdrewritten,roles/forge/README.mdadded.index.htmlanddevbox.jsonleft alone (not deployment). - Left: the comment in
hosts/common/scripts/exec-prestart.shstill says/etc/pushover.env; the script text is part of soft-serve’sExecStartPre, so fix it with the next real change there.
2.8 Optional later: rebuild on the standard layout
infra.rtw.run keeps hardware.nix (root mounted as /dev/sda1 by device
name), the infect-era static networking and network-autodetect.nix. To get
it onto vm.nix + hetzner.nix like every other host, without a reinstall
in place:
- create a new Hetzner VM with the fleet’s
new-host(nixos-anywhere, disko, the stored host key), namedinfra.rtw.runin the inventory - restore
/var/soft/datawithrecover.sh(also a test of the restore), runforge-rebuild-site - stop pushes, run a final
backupon the old box and disable its backup timer (rclone syncdeletes, so a stale box must never sync to R2 again), restore again, move DNS - soft-serve’s SSH keys live in
/var/soft/data/sshand come back with the restore, so clone URLs don’t warn
A good moment for this is when the CI runner is added, if the box needs to be bigger anyway.
Phase 3: moods exports modules; fleet adopts infra.mymood.at
3.1 Refactor moods (in the moods repo)
Each step is verified by just build in moods producing an unchanged
closure for infra.mymood.at.
- No more
fleet/hostspecialArgs in service modules.app.nixreadsfleet // host.settings. Replace with an optionmoods.settings(attrs), defaulting tonix/fleet.nixviamkDefault.secrets.nixuseshost.namefor host-scoped store paths. Useconfig.networking.fqdn(vm.nix sets hostName + domain). Two projects can’t both get a specialArg namedfleeton one host: sovrn’s means mail domain/ACME/relay, moods’ means Jetstream/hostnames.
- Prefix key attributes (
moods-anthropic-api-key, file name unchanged vianame). - Export
nixosModules.services(moods.nix, secrets.nix) andnixosModules.app(app.nix, proxy.nix). Keephosts.nixworking on top of them until the fleet takes over. - proxy.nix sets no global Caddy options; keep it that way (the fleet owns
services.caddy.email).
3.2 Adopt the host
- Add
infra.mymood.attohosts.json(copy from moods’ inventory,roles: ["moods"]). just diff-host infra.mymood.atagainst the running system shows no differences. Differences here come from base/provider modules; fix the fleet copies until it’s clean.colmena apply --on infra.mymood.at(no-op).- Remove
deployfrom moods’ Justfile, or make it print “deploy from ~/projects/servers: just deploy-project moods”.
From now on, shipping moods is:
cd ~/projects/servers
nix flake update moods # picks up moods' last commit (jj: @-)
just deploy-project moods # colmena apply --on @moods
jj commit -m "deploy moods <rev>" # the lock bump is the deploy log
git+file inputs only see committed content.
3.3 Deploy: guard rails for sharing
Before adding anything else to the box:
- Swap: zram (decided 2026-10-06):
zramSwap = { enable = true; memoryPercent = 25; }, about 2 GB of compressed swap in RAM. No disk or layout change, no I/O on netcup’s shared storage, and idle BEAM/Python heap compresses well. It’s a cushion for spikes, not extra capacity; the embedder limits below are what keep the box inside its budget. moods-embedder:MemoryHigh~2.5G,MemoryMax~3G,CPUWeight = 50(default 100). Set these in moods’ module as options with these defaults, or in the fleet as host overrides.- Check after deploy:
systemctl show moods-embedder -p MemoryHigh -p MemoryMax -p CPUWeight; embeddings still keep up (moods’ backlog doesn’t grow).
Phase 3 record (done 2026-10-06)
- 3.1 moods (moods repo):
app.nixreads amoods.settingsoption (defaults fromnix/fleet.nix),secrets.nixusesnetworking.fqdn; no morefleet/hostspecialArgs. ExportsnixosModules.services(options, overlay) andnixosModules.app. Built to the running system; VM test passes. Key attributes and units prefixedmoods-(files unchanged) as a separate commit and deploy. moods’just deploynow refuses and points at the fleet. - 3.2 adoption:
moodsrole inhive.nix(settings fromhosts.jsonsettings.moods), infra.mymood.at inhosts.json(stateVersion26.05). Built tojh2lndcc…(what it ran); deployed: still generation 7, nothing restarted. Host key pinned (it was already). - 3.3 guard rails (
hosts/infra-mymood-at/default.nix; standard hosts now pick uphosts/<name>/when it exists): zram 25% (1.9 GB) andmoods-embedderMemoryHigh=2.5G,MemoryMax=3G,CPUWeight=50. After: no memoryhigh/maxevents, health green, embedding queue near 0. - Prefix deploy: nine key units renamed, moods restarted once, all keys
present,
/healthzok, mymood.at 200.
Phases 4–11: sovrn onto the fleet, Ansible retired (revised 2026-10-06)
sovrn’s own epic (bug 3beadb2, “migrate provisioning from Ansible to NixOS
+ Colmena”) has its first seven children done (Hetzner provisioning, keys,
packages, metrics pilot, Stalwart module, apply plan, orchestration). The
earlier Phases 4–8 here covered only the relay out of what remains; this is
the full path, mapped to the bugs. sovrn has no live hosts (mx99 and
infra.sovrn.at were torn down), so nothing is migrated, only rebuilt.
Decisions behind it (2026-10-06):
- All deployments come from this repo. sovrn’s
Justfile.nix/Justfileand moods’Justfilekeep their dev, test and build recipes; their host recipes (deploy, new-host, known-hosts, check/gen secrets, CI staging) become stubs that calljust -f ../servers/Justfile …or print where things live. - Telemetry edge endpoint is an option. Cells push to
https://infra.sovrn.at(Caddy, basic auth); an edge on the same box as the metrics stack pushes straight to the loopback ingest (http://127.0.0.1:8428VictoriaMetrics,http://127.0.0.1:9428VictoriaLogs), no auth. - The
celllabel is the cell’s hostname on cells and the project name for everything else (moods,sovrn,servers), so the existing alert rules keep working and non-cell services are still grouped. - Rotate the OIDC and OAuth attestation keys first (bug
3418287recorded both leaking to a terminal; nothing uses them yet).
| Phase | What | Bugs |
|---|---|---|
| 4 | Key rotation; sovrn fleet integration; project Justfile stubs | new fleet-integration child of 3beadb2, 3418287 |
| 5 | Relay role | b45a76f (d43eebe closes with it) |
| 6 | Telemetry edge | a6edca3 |
| 7 | Backups | 0e32e5a |
| 8 | Metrics, relay and edge onto the shared box | 6cae3a0 (partly) |
| 9 | mx99 rebuilt by the fleet; live acceptance; CI | 63eb0d9, 6cae3a0, c3b0d04 (part) |
| 10 | Recovery, or an explicit decision to go without it for now | 1badc33 |
| 11 | Delete Ansible; cleanup in sovrn and moods | c3b0d04 |
Not blocking Ansible’s retirement, but tracked: dfb37b1 (sovrnd
/metrics; until it exists BackupStale and TlsExpiring can never fire,
so do it before relying on alerts), edbcac6 (alerts on new forwarding
channels), dc80d76 (Stalwart’s ASN/Geo download), af17761 (nixos-26.11,
which now moves sovrn, moods and the shared box together).
Phase 4: Key rotation, sovrn fleet integration, Justfile stubs
- Rotate
sovrn/shared/oidc-key.pemandsovrn/shared/oauth-attestation.keywith sovrn’s own generator (go run ./cmd/gensecrets --dir <tmp>), into the store, the temp dir removed. Check nothing outside the store pins the old public halves (published JWKS, client metadata). - Options instead of specialArgs.
fleet(nix/fleet.nix) becomessovrn.fleetwithmkDefaults;host.name→config.networking.fqdn;host.pdsOriginSuffix→sovrn.cell.pdsOriginSuffix. Two projects can’t both get a specialArg namedfleet. - Hostnames as options. On the shared box the machine is
infra.mymood.at, but metrics servesinfra.sovrn.atand the relay ismxb.eu.sovrn.at:sovrn.metrics.hostname,sovrn.relay.hostname. Nothing in these roles may use the machine’s own name. - Key names. Prefix key attributes
sovrn-(files unchanged, as moods did); replace the hard-coded"pds-resend-api-key-key.service"and"metrics-password-hash-key.service"inroles/metrics.nixwith a.unitattribute; metrics and relay must both fit on one host. - Caddy email (
roles/metrics.nix,cell-edge.nix):mkDefault, it’s one global value per host. - Export
nixosModules.services,.metrics,.cell(and.relay,.telemetryEdgeas they land), each bringing the overlay. - Roles from the inventory, not the hostname: the fleet’s
hosts.jsonsayssovrn-cell,sovrn-metrics,sovrn-relay. - Host secrets generation moves to the fleet. The per-role generators
(
ssh_host_ed25519_key, Stalwart recovery password, ZDS secrets, …) stay defined by sovrn and are exposed from its flake (an app the fleet’sgen-host-secretscalls);new-hostis merged into the fleet from sovrn’s and moods’ versions. - Stubs. sovrn’s
Justfile.nixhost recipes (new-host,known-hosts,deploy,gen-host-secrets,check-secrets,ci-staging) and moods’ (new-host,known-hosts,deploy,build,eval-hosts,gen-host-secrets,check-secrets) call the fleet’s Justfile or print where to go. - Verified by sovrn’s VM tests (
cell,stalwart,metrics) andstalwart-plan; there is no running sovrn system to compare against.
Phase 4 record (done 2026-10-06).
- Keys rotated: new
sovrn/shared/oidc-key.pemandoauth-attestation.keyfromgensecrets; the leaked copies renamed*.age.leaked-2026-10-06in the store (not under version control; delete when you like). Both public halves are derived by sovrnd at runtime, so nothing else pinned them. The old Ansible vault still holds the leaked values until Phase 11. - sovrn’s uncommitted cell data plane (
63eb0d9, pending review) became its own change first (go test and all four checks passed), so this builds on it and the fleet can see it. - sovrn (
wwzzsxmw):options.nix(sovrn.fleet,sovrn.cell.*,sovrn.metrics.hostname) instead of thefleet/hostspecialArgs;services.nixsplit out ofbase.nix; keys prefixedsovrn-with a.unitattribute; Caddy emailmkDefault; exportsservices,metrics,cell. VM test nodes identical through the options move except tmpfiles rule order; all checks pass after the prefixing. - sovrn (
loyrlypn) and moods (osmtztul): host recipes are stubs that run the fleet’s recipes in its dev shell;fleetexplains the layout. sovrn’sdeploystill refuses uncommitted UI assets (the fleet ships only the last commit). sovrn exportsapps.gen-secret. - Fleet:
sovrn-metricsandsovrn-cellroles (settings fromsettings.sovrn,pdsOriginSuffix);apps.gen-secret-sovrn;just new-host(from moods’, plus roles and stateVersion) andjust gen-host-secrets(DRY=1). infra.mymood.at’s SSH host key linked intoservers/hosts/. mx99 has no stored host key; Phase 9 generates one. - Evaluated, not deployed: moods +
sovrn-metricson infra.mymood.at merges cleanly (keysmoods-*andsovrn-*apart, vhostsmymood.at,feeds.mymood.at,infra.sovrn.at, moods’ unit unchanged); a throwawaysovrn-cellhost evaluates;sovrn-cellbesidemoodsfails the exclusive assertion;DRY=1 gen-host-secretsfor a cell lists the right values. Both live hosts unchanged with the new inputs. - For Phase 8: with sovrn on the shared box, Caddy’s global ACME email
becomes sovrn’s
[email protected](amkDefault) for every site there, mymood.at’s too. Settled in Phase 6:[email protected]on every host.
Phase 5: Relay role (b45a76f)
nix/modules/roles/relay.nix, porting deployment/roles/stalwart_backup
(install, bootstrap, routesync) and ADR-0009 D28 with its 2026-09 amendment:
- Stalwart via
sovrn.services.stalwartwithnix/stalwart/relay.plan.json(arelaycomposition: common + relay; no roles/sovrnd plans): AcmeProvider and a Domain for the relay hostname, SystemSettings, listeners reconciled to:25only plus http on 127.0.0.1:8080, queue expiry 5–7 days,allowRelayingelse = false. Added tochecks.stalwart-plan. sovrn-route-sync(pkgs.sovrn.backup-route-sync) polls each cell’sGET /backup/domainsevery 60 s with the shared bearer token, state in/var/lib/sovrn-route-sync, owning the per-cell pinnedRelayroutes (never plainMx), schedule, strategy and the RCPT relaying exception. The list of cells is a module option. With no cells it converges closed and rejects everything. Decide whether it keeps the recovery admin or gets a restricted account, like sovrnd’s role.- Certificate: Stalwart’s own ACME HTTP-01 for the relay hostname through a
Caddy
http://<relay hostname>vhost that proxies only/.well-known/acme-challenge/to Stalwart (ascell-edge.nixdoes). On the shared box this is just another vhost; the relay doesn’t own 80⁄443. - Firewall: 25 (80 comes with Caddy).
A health check that includes an outbound-25 connect: not needed, the relay delivers to cells on 2525 (see the Phase 5 record).- VM test
checks.relay: relay plus a fake cell; mail for the cell’s domain queues while the cell is down and drains when it returns. This isb45a76f’s done-condition andd43eebe’s, short of live DNS.
Phase 5 record (done 2026-10-06).
- Cells take relayed mail on 2525 (user’s question: can the relay deliver
on an open port instead of 25?). Yes, and it fixes a flaw in the port-25
design. Stalwart 0.16 treats every port but 25 as submission (auth
required; SPF, DMARC and reverse-IP checks off) and on 25 checks
SPF/DMARC against the connecting IP, which for relayed mail is the
relay’s. In the default relaxed mode that doesn’t reject, but legitimate
mail queued during an outage would be recorded as SPF/DMARC failures, fed
to the spam filter, and trigger DMARC failure reports. The cell’s
relaylistener (2525) needs no auth, keeps SPF/DMARC off (the relay checked the real sender on its own port 25), adds a Received header, still runs the spam filter, and is firewalled tosovrn.cell.relaySources. The relay needs no outbound port 25 at all, so netcup’s default SMTP block can’t break it (the earlier “outbound-25 health probe” item is no longer needed). - sovrn commits:
wvrkyoyu(cell listener and overrides, firewall, route-syncCELL_PORT),krpvlntz(two bugs the relay test found: the registry client sentaccountId "a", which only works for callers holdingimpersonate; route-sync’s retry-interval wire shape was rejected by Stalwart),nkprwoko(relay role,relay.plan.jsonwith a restricted route-sync role and account, export,gen-secret). - checks:
relay(new; fake cell down then up: queued, then delivered on 2525 at the first retry),cell(2525 accepts unauthenticated mail and offers no SASL, 587 answers503 5.5.1, firewall rule only for the relay),stalwart-plan(now with the relay composition),stalwart,metrics, Go tests: all pass. - Fleet:
sovrn-relayrole (hostname fromsettings.sovrnRelay, cells = everysovrn-cellhost);sovrn-cellgetsrelaySources= everysovrn-relayhost’s IPv4 and IPv6. Evaluated with infra.mymood.at running moods + metrics + relay and a throwaway cell: route-sync points at the cell on 2525, keys and vhosts don’t collide, moods’ unit is unchanged, the cell’s firewall admits exactly the netcup addresses. - Known gap, unchanged from the Ansible design: bounces for mail that expires on the relay follow the outbound strategy’s fallback (the first cell), which won’t relay for a foreign domain, so they are lost. Fix later by routing non-cell destinations through the outbound smarthost (SMTP2GO on 2525).
- For Phase 8: set
settings.sovrnRelay.hostname = "mxb.eu.sovrn.at"on infra.mymood.at; PTR (IPv4 and IPv6) and DNS formxb.eu.sovrn.at; per-host secrets viajust gen-host-secrets infra.mymood.at.
Phase 6: Telemetry edge (a6edca3)
A sovrn.telemetry.edge module (exported for the fleet) on cells and on the
shared box:
- node_exporter; vmagent scraping it, Stalwart’s
/metrics/prometheusand sovrnd’s/metricswhere present, remote-writing to the endpoint; Vector shipping journald to VictoriaLogs. - Options:
endpoint.metrics/endpoint.logs(default theinfra.sovrn.atURLs with basic auth frommetrics-password; on the metrics box the loopback ingest without auth),cell(default the host name), and a unit/job → project map so that on the shared box moods’ units carrycell="moods", sovrn’scell="sovrn", the fleet’s owncell="servers". - Labels must match the existing rules (
cell,service,levelperroles/metrics.nix). - VM test: an edge plus the metrics role; a scraped series and a journald line arrive with the right labels.
Phase 6 record (done 2026-10-07).
- sovrn (
tnqrnrtk):nix/modules/telemetry-edge.nix(sovrn.telemetry.edge, part ofnixosModules.services; also exported asnixosModules.telemetryEdgefor hosts with no sovrn role). node_exporter (127.0.0.1:9100), vmagent (127.0.0.1:8429, 2 GiB disk buffer), Vector (journald → VictoriaLogs’ Elasticsearch ingest, 512 MiB disk buffer), all on the nixpkgs modules. Both daemons are DynamicUsers: the password comes byLoadCredential, Vector reads every journal (ZDS user units included) through thesystemd-journalgroup. No root daemons, unlike the Ansible role. - Endpoints: cells push to
https://<sovrn.metrics.hostname>with basic auth (usercell,metrics-password, now declared by the edge rather than the cell role); the metrics role points its own edge at127.0.0.1:8428/9428without auth. - Labels:
cell,hostandserviceon every series and log line,levelon logs.cell(unchanged contract, alert rules group by it) is who owns the service,host(new, a VictoriaLogs stream field) where it runs; both are the machine’s FQDN by default, so cells are told apart by hostname. A sovrn role on a shared host labels its own services with its own hostname, bothcellandhost(user, 2026-10-07: the relay’s Stalwart must be told apart from the cells’ and from other relays, andmxb.eu.sovrn.atmatters more than the machine’s name): the relay’s Stalwart and route-sync aremxb.eu.sovrn.at, the metrics stackinfra.sovrn.at. Other projects’ units are mapped by prefix (units.<prefix>, longest wins). Set per scrape job (no-remoteWrite.label). Logserviceprefers_SYSTEMD_USER_UNIT, so ZDS instances show aszds@<domain>.serviceinstead ofuser@<uid>.service. - Scrape jobs:
host,vmagent-selfalways;stalwart(/metrics/prometheus) andsovrnd(/metrics,up == 0untildfb37b1) when those services run. Cell, relay and metrics roles enable the edge. checks.telemetry(new): the metrics role labelled like the shared box, plus a cell pushing through Caddy with basic auth. Host series withcell/host/service/mountpoint="/",up{job="vmagent-self"}, a cell log line withlevel="error", amoods-*unit undercell="moods"on the machine’shost, VictoriaMetrics undercell=host="infra.sovrn.at", an unmapped unit undercell="servers", edge ports loopback only.checks.relaynow also runs the metrics role (the shared box’s shape): the relay’s Stalwart logs, route-sync logs andup{job="stalwart"}arrive ascell=host=<relay hostname>, node metrics under the machine’s name.metrics,relay,cell,stalwart,stalwart-planpass (cellonce hit a Stalwart/Pebble challenge race and passed on rerun).- Fleet: a host with
sovrn-metricsorsovrn-relaygetscell = "servers"and, when it runs moods,units.moods.cell = "moods". Evaluated and built with infra.mymood.at as moods + metrics + relay and a throwaway cell: the shared box pushes to loopback without ametrics-passwordkey; stalwart and route-sync aremxb.eu.sovrn.at/mxb.eu.sovrn.at, the metrics stackinfra.sovrn.at/infra.sovrn.at, moodsmoods/infra.mymood.at, host and vmagentservers/infra.mymood.at; everything on the cell ismx99.eu.sovrn.at, pushed tohttps://infra.sovrn.atascell. Both live hosts unchanged by the sovrn bump. - Caddy’s ACME email (user, 2026-10-07):
[email protected]for every project. The fleet’smodules/caddy.nixsets it on every host (live hosts: only the Caddyfile’s globalemailline; not yet deployed); sovrn’sfleet.acme.caddyEmailmakes it sovrn’s default too. Stalwart’s own ACME account keeps[email protected](fleet.acme.contact). This settles the Phase 4 note about mymood.at inheriting sovrn’s address.
Phase 7: Backups (0e32e5a)
Litestream (stalwart.db, sovrn.db, oauth.db, ZDS DBs) and rclone (ZDS blobs
every 15 min, cell config) to the cell’s R2 bucket, credentials as keys,
health markers under /var/lib/sovrn/health with the ages sovrnd’s
/healthz expects (15 m / 45 m). Done on a cell: markers fresh, R2 current.
Phase 7 record (done 2026-10-07).
- sovrn (
vkmnmqtk):nix/modules/backup.nix(sovrn.backup, enabled by the cell role).sovrn-litestreamreplicatesstalwart.db, sovrnd’ssovrn.dbandoauth.db, on the landing cell alsowaitlist.db, and every ZDS database (directory watch) tos3://<bucket>/litestream{,-zds}.sovrn-litestream-freshness(5 min) runslitestream sync -waitper database and refreshes/var/lib/sovrn/health/litestream.ok;sovrn-backup-sync(15 min) rclones Caddy’s storage and the ZDS blobs and refreshesrclone.ok. sovrnd’shealth.*markerfilepoint at them (max ages 15 m / 45 m are sovrnd’s defaults). - nixpkgs Litestream 0.5.17 (Ansible pinned 0.5.13). Snapshot interval
(24 h) and retention (30 days) are now in the top-level
snapshotblock: 0.5 has no per-replicaretention/snapshot-intervaland its YAML parsing ignores unknown keys, so Ansible’s settings were silently dropped. Litestream sets WAL mode itself. - Not copied any more: Stalwart’s
config.json(declarative, in the store; certificates live instalwart.db). Root, as in Ansible (Litestream writes beside databases owned bystalwartandsovrn); R2 keys are credentials passed to the tools as environment, rclone has no config file. Bucket<region>-<cell>from the hostname, endpoint fromsovrn.fleet.r2(account, jurisdictioneu): mx99 getseu-mx99on the EU endpoint, as before. - Guard (new):
rclone syncdeletes and Litestream would take over the replica, so a rebuilt cell must not back up before its data is restored. Every backup unit first runssovrn-backup-guard: it passes once/var/lib/sovrn-backup/armedexists; otherwise it arms the box only if the bucket is empty and refuses (retrying every minute, logged at error) if not. Recovery (Phase 10) arms the box after restoring; to start over on purpose,touch /var/lib/sovrn-backup/armed. checks.backup(new, MinIO as R2, permitted as insecure on the test node only): arming on an empty bucket; every database replicated; a ZDS database created later picked up; the freshness marker; a replica restores with its latest rows; rclone objects and marker; an unarmed box refuses, then runs once armed.cellpasses, its/healthznow reporting the backup legs (failing without R2, as intended).- Fleet: sovrn input bumped; both live hosts unchanged; a throwaway mx99
evaluates and builds with
sovrn-r2-access-key/sovrn-r2-secret-keykeys (present in the store). - Landing cell (user, 2026-10-07): exactly one cell serves
sovrn.at(the landing page; no way to share its certificate across cells): sovrn’sfleet.rootHost,mx99.eu.sovrn.at. Only it gets thesovrn.atCaddy site andlanding.enabled, as under Ansible. Signups land in itswaitlist.db; sovrnd creates an empty one on every cell (it opens it unconditionally), but only the landing cell backs it up. Ansible never backed it up at all. Moving the landing page to another cell leaves the signups in the old cell’s database and bucket: copy them by hand. - Fleet: every
sovrn-cellhost asserts that sovrn’srootHostis asovrn-cellinhosts.json(Ansible’s “landing-root singularity” guard, lost in the port) and that no host overridesrootHostinsettings.sovrn(cells must agree on one). Checked with throwaway cells: mx99 + mx98 build, mx99 backs upwaitlist.dband mx98 doesn’t; mx98 alone and an override on mx98 each fail with their message. - For Phase 9:
eu-mx99held the Ansible mx99’s backups (31 objects, 3.8 MiB); the user deleted them on 2026-10-07, so the guard will arm.
Phase 8: Metrics, relay and edge on the shared box
Prerequisites on the box are in place (Phase 0: headroom, outbound 25 unblocked; Phase 3: guard rails).
- Metrics: DNS
infra.sovrn.at→ infra.mymood.at; addsovrn-metrics; secrets (metrics-password-hash,pds-resend-api-key) are in the store;dry-activate(moods must not restart), deploy. Check:https://infra.sovrn.at/healthz200 and the Larm monitor pointed at it; vmui and logs UI behind basic auth; a test alert reaches[email protected]through Resend; the box still within its Phase 0 budget. - Edge: on infra.mymood.at with the loopback endpoint; moods’ and
sovrn’s units show up under
cell="moods"/cell="sovrn". - Relay: netcup PTR (IPv4 and IPv6) →
mxb.eu.sovrn.at; DNSmxb.eu.sovrn.at→ the box; per-host secrets viagen-host-secrets; addsovrn-relay, deploy. Check: from outsidenc -v mxb.eu.sovrn.at 25gets Stalwart’s banner (the inbound-25 test netcup’s policy might still block); STARTTLS presents a valid certificate; RCPT for any domain is rejected (closed, no cells yet); route-sync healthy. - Leave
backupMx = ""in sovrn’s fleet settings until Phase 9.
Phase 8 record (done 2026-10-07). Four deploys to infra.mymood.at, one change each; moods was never restarted.
- Prerequisites:
infra.sovrn.atandmxb.eu.sovrn.atA/AAAA → the box; PTR for both addresses →mxb.eu.sovrn.at(checked over DNS-over-HTTPS: plain port 53 to public resolvers is blocked from here and from the box). - Generation 12: Caddy’s ACME email
[email protected](the Phase 6 change). Caddy reloaded. infra.rtw.run still runs without it. - Generation 13 (8.1, 8.2):
sovrn-metrics, which brings its telemetry edge. Caddy restarted (new credential); everything else new.infra.sovrn.at: Let’s Encrypt certificate,/healthz200 and public,/vmui/,/select/vmui/,/vmalert/401 without and 200 withmetrics-password; a test alert went out through Resend (alertmanager_notifications_totalemail 1, failed 0). Labels: host metricsservers/infra.mymood.at, moods’ unitsmoods, the metrics stackinfra.sovrn.at. The closure grew by ~860 MiB: nixpkgs’ Vector 0.58 keeps a runtime reference to gcc-wrapper (disk only; worth reporting upstream). - Generation 14: sovrn fix (
qtqttunr): Caddy proxied/am*to Alertmanager unchanged, which served from/, so its UI was a 404 (as under Ansible). Alertmanager now serves under/am(itswebExternalUrl); vmalert’s notifier and/healthzuse the prefix;checks.metricscovers both. - Generation 15 (8.3):
sovrn-relaywithsettings.sovrnRelay.hostname = "mxb.eu.sovrn.at"; per-host secrets fromgen-host-secrets. Stalwart got its Let’s Encrypt certificate through Caddy’shttp://mxb.eu.sovrn.atchallenge proxy; STARTTLS verifies on IPv4 and IPv6; RCPT for any domain gets550 5.1.2 Relay not allowed(no cells); route-sync not running (no cells); telemetrycell=host="mxb.eu.sovrn.at". Memory available after all three: ~5 GiB of 7.75. - Inbound port 25 first timed out from outside (check-host.net, four
nodes) while 443 answered; the box’s firewall accepted 25 and tcpdump saw
no SYN arrive. The user’s netcup firewall policy allowed inbound 22/80/443
only (set up for moods); they added inbound TCP 25 and UDP 443 (HTTP/3).
After that, 25 answers from Germany, Spain and Japan by name, and
infra.sovrn.atanswers over HTTP/3. IPv6 port 25 isn’t checked from outside (check-host.net takes no IPv6 literals). - To do outside the repo: point the Larm monitor at
https://infra.sovrn.at/healthz; confirm the test alert (and its resolution) reached[email protected]. - Noted: Stalwart warns that the box’s resolver (netcup’s, via systemd-resolved) can’t validate DNSSEC, so it disables DANE for outbound delivery. The relay only delivers to cells; revisit with the cells’ resolver settings.
backupMxstays""until Phase 9.
Before Phase 9 (2026-10-07).
- Stalwart’s admin web UI is off (sovrn
lxpptkst, user’s call: sovrnd fronts everything, Stalwart is managed over its API). The plan disables the web UI Application rather than deleting it: with none, Stalwart re-seeds it enabled and fetched from GitHub’slatest. A disabled application is never fetched or served (checked in 0.16.23’s source and inchecks.stalwart). This also drops the only npm build (no binary-cache entry on any architecture). Not yet deployed to infra.mymood.at: the relay’s Stalwart restarts to converge the plan. - arm64 cells (Hetzner has few amd64 servers left): an
aarch64-linuxmx99 evaluates and needs only sovrn’s Go commands and ZDS (Zig) compiled; Stalwart, RocksDB, Vector, Caddy, Litestream, rclone and the Go toolchain come from cache.nixos.org. The disk layout and GRUB already handle UEFI (Hetzner arm64). This machine can’t build aarch64 (qemu binfmt is registered but Nix has noextra-platforms), so either build on the target (ColmenabuildOnTarget, nixos-anywhere--build-on remote) or enable emulation here. sovrn’s VM tests are x86-only, so mx99 is the first run of sovrnd and ZDS on arm64. - RocksDB stays. Stalwart only ever uses SQLite here (
config.jsonnames it before the first start, so bootstrap mode never runs), but nixpkgs builds Stalwart with therocksfeature and the binary linkslibrocksdbdynamically. Dropping it means a custom Stalwart build without cache.nixos.org; keeping it costs ~117 MiB of disk. - vlpds later, as a container (bug
a4a3137, user, 2026-10-07): not before upstream settles (very fast-moving; revisit in about a week, replace ZDS before beta only if it has matured). Phase 9 builds and deploys ZDS. When it comes: upstream’s multi-arch imageghcr.io/jazware/vlpds(not the BTBurke fork, which is being retired), run byvirtualisation.oci-containerswith podman, pinned by digest throughdockerTools.pullImage, so no arm64 Rust build is ever needed. The fleet then compiles only sovrn’s Go commands (and ZDS until then) for a cell: on the cell or under emulation. eu-mx99emptied by the user (checked: 0 objects), so the rebuilt mx99’s backup guard arms on its first run.checks.cellis flaky: Stalwart re-posts an ACME challenge Pebble is still processing (Cannot update challenge with status processing) and then backs off past the 300 s wait; 2 failures in ~6 runs, each passing on rerun. Harden it (e.g.PEBBLE_VA_NOSLEEP=1, or a longer wait).
Phase 9: mx99 on the fleet; live acceptance
new-host(fleet) builds mx99 fresh:roles: ["sovrn-cell"], exclusive.- Live checks the epic has been waiting for:
- the domain signup saga end to end on staging (
63eb0d9) - mx99’s metrics and logs in VictoriaMetrics/Logs with labels the alert
rules match (
a6edca3); a staged failure fires the matching alert (6cae3a0) - backup markers fresh, R2 current (
0e32e5a) - the relay pulls mx99’s domains; set
backupMx = "mxb.eu.sovrn.at"; with mx99’s Stalwart stopped, mail queues on the relay and drains when it returns (b45a76f,d43eebe;docs/runbooks/relay-drain-verify.md)
- the domain signup saga end to end on staging (
- CI: sovrn’s
ci-staging(tests, then deploy mx99) runs from this repo: bump the sovrn input, run sovrn’s checks,colmena apply --on mx99.eu.sovrn.at. Needs the CI remote URLs (open decision 1).
Phase 9 record (2026-10-07/08; CI step open).
- mx99 rebuilt as an arm64 staging cell (Hetzner CAX11: 2 vCPU, 4 GB) by
just new-host 2.29.20.216 mx99.eu.sovrn.at hetzner sovrn-cell(pdsOriginSuffix = ".pds1.eu.sovrn.at"pre-set). New host key; R2 and ZDS secrets reused from the store,sovrnd-stalwart-passwordgenerated. The firstjust deployuploads the Colmena keys nixos-anywhere doesn’t. DNS: A/AAAA, PTR,sovrn.at(landing), wildcard*.pds1.eu.sovrn.at. - Building for arm64: this machine builds aarch64 under qemu (user
enabled
extra-platforms): Go commands 729 s, ZDS 2105 s. mx99 is also a Nix remote builder (/etc/nix/machines, with its host key): Go commands 219 s cold / 146 s warm, and production cells reuse the same store paths. The user is fine with the staging cell spending resources on builds. - Live: Stalwart converged (17 created), LE certificate; backup guard armed on
the emptied
eu-mx99, Litestream and rclone markers fresh;/healthzok; metrics and logs at infra.sovrn.at undercell="mx99.eu.sovrn.at"; all four scrape jobs up. (Correction, 2026-10-08: sovrnd has no/metrics; it redirects to/login, and vmagent reports the jobupwith 0 samples. Tracked in sovrn dfb37b1.) - Alert drill:
CellErrorBurstfired for mx99 and was emailed. (The first drill didn’t run:systemd-rununits have noseq/sleepon PATH; rerun with--setenv=PATH=/run/current-system/sw/bin.)CellDiskFullSoonfired falsely on mx99 from install/build bursts: staging cells (fleet.stagingCells) now judge the trend over 6 h held 1 h (sovrnafba45d, infra.mymood.at generation 18). - Relay covers mx99 (infra.mymood.at generation 17, route-sync started);
Stalwart’s web UI off on the relay (generation 16);
backupMx = "mxb.eu.sovrn.at"deployed to mx99. - sovrnd fixes found live: a first login was refused when the user had no
service record (
c2db915); domain activation rolled back because the PDS health gate didn’t wait for ZDS to listen (6efaf8b); the freshness probe raced Litestream’s socket on first boot (2883fe7). Each has a test. - Domain signup (
test.kilimanjaro.io, user-created through sovrnd): ZDS instance active, DNS checks passed, mailboxtest@. - Mail: outbound via SMTP2GO passes DKIM (s931828), SPF (em931828
return path) and DMARC at Fastmail; inbound from Gmail delivered (Fastmail
can’t test it: it hosts
kilimanjaro.ioand keeps subdomains internal). - Relay queue-and-drain passed: Stalwart stopped on mx99; Gmail fell back
to
mxb.eu.sovrn.at; queued; first attempt to mx99:2525 refused; Stalwart restarted; delivered on 2525 with STARTTLS at the next retry (~5 min); mailbox ingested it. - Follow-ups filed:
9cbbce8(cells must not claimsovrn.at; internal domain for service accounts),2b0ddfc(Apple configuration profile),0444db1(guided device setup). iOS needed SMTP credentials entered separately and a stale account re-added. - Noticed: Stalwart’s inbound spam filter takes ~30–60 s per message on both mx99 and the relay (likely DNS-based list lookups timing out); doesn’t block delivery but worth a look. Stalwart logs DANE disabled (resolvers don’t validate DNSSEC).
- Open: CI (sovrn checks, then
colmena apply --on mx99, open decision 1).
Phase 10: Recovery (1badc33)
sovrn-restore.target and sovrn-freshen-backups as in the bug, rehearsed
on a scratch cell. The Ansible playbooks are the only restore path today, so
either this lands before Phase 11, or decide explicitly to delete Ansible
without a restore path while there is no production data.
Phase 10 record (done 2026-10-08).
- Design (user): restore is automatic and keyed off the backup guard, with no install flag. Cells keep their primary IPv4 and IPv6: Hetzner Primary IPs move to the replacement, so DNS and PTR never change.
- sovrn
c6bfc5b:sovrn-restore(oneshot, before and required by Stalwart, sovrnd, Caddy and the backup units; waits for its R2 keys, retries until R2 answers). Armed or local databases present: nothing. Empty bucket: a new cell, armed. Otherwise: restore every database from Litestream (Stalwart, sovrnd’s,waitlist.dbon the landing cell, every ZDS database),PRAGMA integrity_check(any failure stops the unit, so no service starts), owners, ZDS blobs and Caddy storage copied back (never synced), arm.sovrn-freshen-backupsstops the services and forces the final Litestream sync and rclone run on a cell about to be replaced.checks.restore(new) covers new cell, backup, freshen, empty replacement restore, continued backups, and never overwriting local data. - Rehearsal on mx99: 5 messages in
[email protected]; freshen; old CAX11 powered off; Primary IPs moved to a new CAX11; a Gmail message during the outage queued on the relay (route-sync kept the last-known domains);just new-hostandjust deploy;sovrn-restorerestored all 5 databases before any service started; same accounts and domains; no new ACME orders (Stalwart and Caddy certificates came from the backup); ZDS back with the same DID and repo revision;/healthzok; the relay delivered the outage message at its next retry; the phone shows all 6 messages. - Bug found and fixed (sovrn
9c9ee66): sovrnd didn’t wait for thezds-*key units (Colmena uploads sovrn-owned keys after activation), so on first boot it skipped the ZDS env and left the restored instance down until restarted. Deployed. - The relay’s retry backoff grows (1, 5, 15 minutes), so mail sent during an outage can arrive up to ~15 minutes after the cell is back.
- Left for Phase 11: rewrite
docs/runbooks/recovery-primary-ip.mdandpds-cell-recovery.mdfor this flow; delete the old mx99 server once the user is satisfied.
Phase 11: Decommission Ansible and clean up (c3b0d04)
- sovrn: delete
deployment/(roles, playbooks, inventory, vault, scripts),nix/pkgs/stalwart.nix+ its Cargo.lock,bootstrap-stalwart.sh, the ansible package in devenv;nix/hosts.nix,nix/hosts.json,nix/modules/{base,hetzner}.nix; foldJustfile.nixintoJustfile(stubs from Phase 4 stay). Rewritedocs/deployment.mdand the runbooks for the fleet. Close3418287(superseded),d43eebe,6cae3a0. - moods: delete
nix/hosts.nix,nix/hosts.json,nix/modules/{base,vm,hetzner,netcup}.nix(the fleet has them). - Both keep
fleet.nixas their settings defaults andupdate-nixpkgs; bumping nixpkgs means sovrn, moods and the fleet together (just sync-nixpkgs). - Optional: one small library flake for the three near-identical secrets modules.
Phase 11 record (done 2026-10-08).
- sovrn (
tsxnqxwz):deployment/deleted (Ansible roles, playbooks, inventory and the encrypted vaults; 77 files). Also gone:nix/hosts.nix,nix/hosts.json,nix/modules/hetzner.nix,nix/keys/, thecolmenaHive/nixosConfigurationsoutputs and thecolmenaanddiskoinputs, colmena and nixos-anywhere from devenv.nix/modules/base.nixbecamenix/tests/node.nix(the VM test nodes’ base only).Justfile.nixfolded intoJustfile(dev loop, gen, tests, the fleet stubs;ci-stagingdropped). Docs rewritten for the fleet:docs/deployment.md, runbooksprovision-cell.md,recovery-primary-ip.md(as rehearsed in Phase 10),pds-cell-recovery.md,metrics-box.md,relay-drain-verify.md,nix/README.md; secret-provisioning wording in the design docs. All VM checks pass. - moods (
ltsvqxsy):nix/hosts.nix,nix/hosts.json,nix/modules/{base,vm,hetzner,netcup}.nix,nix/keys/, the host outputs and inputs, colmena and nixos-anywhere from devenv. Its VM test passes. - CI: none for now (user, 2026-10-08): tests and deploys run from this machine; staging (mx99) first, production cells after. Phase 9’s CI step and open decision 1 are dropped until that changes.
- Bugs closed (sovrn):
b45a76f,d43eebe,63eb0d9,0e32e5a,a6edca3,1badc33,ae42c89,6cae3a0,ebce532,3418287(superseded),c3b0d04, and the epic3beadb2. Still open by design:dfb37b1(sovrnd lacks the backup-marker and TLS-expiry series, soBackupStale/TlsExpiringcan’t fire),af17761(nixos-26.11),edbcac6,dc80d76, and the follow-ups filed during Phases 9–10 (9cbbce8,2b0ddfc,0444db1,d4988da,6844d4f,a6c4223). - The old Ansible vaults (with the keys rotated in Phase 4) remain in sovrn’s history; rewrite it if that matters.
Rollback summary
| Step | Rollback |
|---|---|
| 2.3, 3.2 | Nothing changed on the host |
| 2.4, 2.6 | Previous generation (GRUB or nixos-rebuild switch --rollback); agenix ciphertext in git until 2.7 |
| 2.5 | Previous generation; soft-serve data from the step-2 R2 backup if its database migrated |
| 3.3, 8 | Remove the role from hosts.json and deploy; DNS for infra.sovrn.at / mxb.eu.sovrn.at can stay |
| 4 (keys) | None needed: nothing uses the keys yet |
| 9 | mx99 is staging: rebuild it |
Risks and gotchas
- Two hives, one host. Once a host is in the fleet, deploying it from a project repo would erase the other projects’ services. Remove the projects’ deploy recipes as each host is adopted.
- Root key lockout. Resolved for infra.rtw.run: the hetzner key is in the
config since deploy 2 (generation 69) and logs in on its own. When testing
which key works, use
ssh -F /dev/null …: aHostblock’sIdentityFileis offered even alongside-iandIdentitiesOnly=yes. - netcup’s firewall policy. The outbound SMTP block lived in netcup’s policy, not on the box. If it comes back, the relay accepts mail it can’t deliver. The relay health check probes outbound 25 (Phase 5).
- The relay is a public SMTP listener on a box with other secrets. Each
service’s keys are root-owned and passed by
LoadCredential; Stalwart and moods run as separate sandboxed users. Keep it that way: no shared service users, no group-readable key directories. - Shared nixpkgs is a coordination point. sovrn needs unstable for Stalwart 0.16; moods rides along, and so do metrics and the relay. When nixos-26.11 ships, sovrn’s plan is to move to it, and the shared box moves with it. infra.rtw.run stays on its own stable channel.
- Deploying one project redeploys the host. A moods deploy rebuilds the shared box at the fleet’s pinned sovrn version too. Units whose config didn’t change aren’t restarted, but a nixpkgs bump restarts most things, including the relay’s Stalwart (senders retry; the relay is the backup).
- pgit is hosted on the forge. The
pgitinput isgit+https://git.kilimanjaro.io/pgit. With the forge down, a fresh evaluation can’t fetch it. Mirror it elsewhere (or keep it in the local store) before relying on the fleet for recovery. rclone syncdeletes. Never let a box with empty or stale/var/soft/datarun the backup timer (2.8).- Caddy vhost merging is silent.
extraConfigis alinesoption, so the same vhost declared twice is concatenated, not rejected.modules/caddy.nixshould assert vhost names are unique per project. stateVersionis per host. infra.rtw.run 25.11, infra.mymood.at 26.05. Never take it from a shared base.
Open decisions
Where CI runs: none for now (2026-10-08); everything runs from the operator’s machine.buildOnTargetfor infra.rtw.run.- Whether and when to rebuild infra.rtw.run on the standard layout (2.8), e.g. together with adding the CI runner.
Decided:
- All deployments come from this repo; project Justfiles keep dev/test recipes and stub out host recipes (Phase 4).
- Telemetry edge: endpoint is an option, loopback on the metrics box;
celllabel = cell hostname, else the project name (Phase 6). - Rotate sovrn’s OIDC and OAuth attestation keys first in Phase 4.
- Swap on the shared box: zram (3.3).
- Relay hostname:
mxb.eu.sovrn.at. The netcup box is in us-east; it backs up the EU cells from a different cloud and region on purpose, for redundancy. The name follows the region it serves, not where it runs.