fleet-migration-plan.md

Fleet migration: sovrn, moods and servers → one Colmena fleet

Status: Migration complete (2026-10-08). Every host is deployed by this repository from the operator’s machine: infra.rtw.run (forge), infra.mymood.at (moods, sovrn’s metrics stack infra.sovrn.at and backup relay mxb.eu.sovrn.at), mx99.eu.sovrn.at (arm64 staging cell). sovrn’s Ansible is gone; recovery is automatic and rehearsed.

Goals

  1. infra.rtw.run on the standard method. Move it off agenix, the Makefile and nixos-rebuild onto Colmena and the secrets tool, and bring its packages up to date (it hasn’t been updated since June). It stays a separate box: it’s the base infrastructure for every project (personal git hosting now, a CI runner later).
  2. Consolidate onto the netcup box. infra.mymood.at runs moods today. Add sovrn’s metrics box and sovrn’s backup mail relay to it instead of building two more servers.

One Colmena hive, living in this repo, deploys every host. sovrn and moods keep developing independently (their own flakes, VM tests, devenv), and the fleet composes them onto hosts.

Principles

  1. One hive owns each host. Colmena deploys a whole system; two hives pointed at one box would each erase the other’s services. Once a host is in the fleet, only the fleet deploys it.
  2. Projects export, the fleet composes. sovrn and moods export overlays and NixOS modules; they never import the fleet. The fleet pins each project in flake.lock, so “what is deployed” is a commit in this repo.
  3. Nothing is updated implicitly. No nix flake update as a side effect of deploying. Bumping a project or nixpkgs is its own commit.
  4. Prove no-ops before changing behaviour. Every adoption and refactor first produces a system closure identical to what’s running (checked with nvd diff), then a separate deploy makes the actual change.
  5. One behavioural change per deploy, each with a rollback (previous generation, or colmena apply from the previous commit).

Phase 0 findings (2026-10-06)

Live hosts

Only two hosts are live. sovrn’s Ansible hosts (mx99.eu.sovrn.at, the staging cell, and infra.sovrn.at, the metrics box) were torn down during sovrn’s move to Colmena. They get rebuilt from scratch on the fleet, so there is no data to migrate.

Host Provider Size Runs OS
infra.rtw.run Hetzner 2 vCPU, 3.7 GiB, 38 GB soft-serve 0.8.5, Caddy 2.10.0, pgit NixOS 25.05 (ac62194, 2026-01-02), generation 68, up 114 days
infra.mymood.at netcup 4 vCPU, 7.75 GiB, no swap, 125 GB moods (service + embedder), Caddy NixOS 26.11pre (7a0f122), deployed by moods’ hive

infra.rtw.run

infra.mymood.at: headroom

Measured:

RAM CPU Disk
moods-embedder 1.75 GB, flat (peak = current) ~1.4 cores average (447k CPU-s over 3.8 days) model in the Nix store
moods (BEAM) 1.1 GB cgroup, includes SQLite page cache ~0.03 cores /var/lib/moods 3.6 GB
Caddy + system ~0.2 GB – –
Box 2.35 GB used, 5.6 GB available load ~1.1, CPU pressure ~1.4% 5.9 / 125 GB

Estimated additions (nothing to measure yet):

RAM CPU Disk
sovrn metrics (VictoriaMetrics, VictoriaLogs, 2× vmalert, Alertmanager, healthz), 1–3 cells 0.5–1 GB <0.2 core VictoriaLogs capped at 25 GiB; VictoriaMetrics 90d, a few GB
sovrn relay (Stalwart store-and-forward + route-sync) 0.2–0.4 GB ~0 idle queue, small
Projected total ~3.5–4 of 7.75 GB ~1.7 of 4 cores ~35–40 of 125 GB worst case

It fits, with two guard rails (Phase 3.3): swap (there is none, so a spike goes straight to the OOM killer) and limits on moods-embedder (MemoryHigh/MemoryMax, lower CPUWeight) so an embedding backlog can’t starve Stalwart or Alertmanager.

infra.mymood.at: is co-locating safe?

infra.mymood.at: network

Where things are today

sovrn moods servers
Deploy Colmena hive (Justfile.nix), no hosts; Ansible in deployment/ (hosts torn down) Colmena hive (Justfile) Makefile → scripts/update.sh → nixos-rebuild on the target
NixOS hosts none (nix/hosts.json is {}) infra.mymood.at infra.rtw.run
nixpkgs unstable 7a0f122 (Stalwart 0.16) unstable 7a0f122 (kept equal to sovrn) running 25.05 ac62194; repo HEAD says 26.05
Secrets secrets CLI → Colmena keys in /var/lib/sovrn-keys same, /var/lib/moods-keys agenix, encrypted to the host SSH key
Install nixos-anywhere + disko, host key from the store same, plus netcup cloud-init + nixos-infect, hardware.nix fetched afterwards
Caddy services.caddy services.caddy hand-written systemd.services.caddy + Caddyfile
Relay nix/modules/roles/relay.nix is a placeholder; Ansible role deployment/roles/stalwart_backup (install, bootstrap, routesync) to port from – –

Target

Host Roles Notes
infra.rtw.run (Hetzner) forge (later ci) Separate base infrastructure. In place on its legacy layout first; optional rebuild later (2.8)
infra.mymood.at (netcup) moods, sovrn-metrics, sovrn-relay The shared box. DNS names infra.sovrn.at and mxb.eu.sovrn.at point here
sovrn cells (Hetzner, new) sovrn-cell Exclusive. mx99 rebuilt first as staging

Cost: metrics and the relay each needed their own box in the Ansible design. Here they cost nothing extra; the fleet is two always-on boxes plus however many cells sovrn runs.

Shared box port map

Owner Ports
moods 7071 (embedder), 7072 (service), loopback
sovrn metrics 8428, 9428, 8880, 8881, 9093, 8090, loopback
sovrn relay 25 public; 8080 (Stalwart http), 18080 (recovery mode), loopback
Caddy (shared) 80, 443/tcp, 443/udp

Caddy vhosts on the shared box: mymood.at, feeds.mymood.at (moods), infra.sovrn.at (metrics), http://mxb.eu.sovrn.at (ACME challenge proxy to Stalwart only).

Repository layout (this repo)

servers/
  flake.nix              inputs: nixpkgs (= sovrn/moods rev), nixpkgs-stable,
                         colmena v0.5.0, disko, pgit, sovrn, moods
  flake.lock             the deployed versions of everything
  Justfile               eval, build, diff-host, deploy, deploy-project,
                         check-secrets, known-hosts, sync-nixpkgs
                         (new-host, gen-host-secrets in Phase 4)
  hosts.json             inventory (below)
  hive.nix               inventory → Colmena hive; also nixosConfigurations;
                         fails evaluation if sovrn's or moods' nixpkgs differs
  keys/admins.pub        root SSH keys for standard hosts (same as sovrn/moods)
  modules/
    base.nix             every standard host: SSH, firewall, resolved, nix GC, admin keys
    vm.nix               disko + GRUB + networkd   (from moods)
    hetzner.nix          from moods
    netcup.nix           from moods
    secrets.nix          `servers.secrets`, /var/lib/servers-keys   (moods' module, renamed)
    fleet.nix            `fleet.roles`; assertion: an exclusive role (sovrn-cell) is alone on its host
    caddy.nix            fleet-wide Caddy settings (ACME email); vhost-collision assertion still to come
  roles/
    forge.nix            soft-serve, pgit hook, backup timer, Caddy vhosts
  hosts/
    infra-rtw-run/       legacy layout: hardware.nix, static networking
  docs/

Inventory (hosts.json)

moods’ schema plus roles, nixpkgs and layout:

{
  "infra.mymood.at": {
    "provider": "netcup", "system": "x86_64-linux", "disk": "/dev/vda",
    "ipv4": "152.53.243.113", "ipv4Prefix": 22, "ipv4Gateway": "152.53.240.1",
    "ipv6": "2a0a:4cc0:2000:84c4:c4a5:5fff:fe80:c760/64",
    "sshHostKey": "ssh-ed25519 …",
    "roles": ["moods", "sovrn-metrics", "sovrn-relay"],
    "settings": { "moods": {}, "sovrn": {} }
  },
  "infra.rtw.run": {
    "layout": "legacy", "nixpkgs": "stable", "system": "x86_64-linux",
    "ipv4": "178.104.201.207", "ipv6": "2a01:4f8:1c18:235d::1/64",
    "sshHostKey": "ssh-ed25519 …  (from identities/host/infra-rtw-run.pub)",
    "roles": ["forge"]
  }
}

What sovrn and moods export

Output sovrn moods
overlays.default exists exists
nixosModules.services stalwart, sovrnd, zds, secrets, sovrn user (all off by default) moods.nix, secrets.nix
nixosModules.<role> metrics, relay, cell app (app.nix + proxy.nix)
checks VM tests (stay; add relay) VM test (stays)

Removed from the projects at the end (Phase 11): hosts.nix, hosts.json, base.nix, provider modules, new-host/known-hosts/deploy recipes.

Secrets

One secrets store (SECRETS_DIR, ~/projects/data), one prefix per project, one key directory per project on the host:

Namespace Store paths On host
sovrn.secrets sovrn/shared/…, sovrn/hosts/<fqdn>/… /var/lib/sovrn-keys
moods.secrets moods/shared/…, moods/hosts/<fqdn>/… /var/lib/moods-keys
servers.secrets servers/shared/…, servers/hosts/<fqdn>/… /var/lib/servers-keys

Colmena’s deployment.keys is one attrset per host, and each key gets a <attr>-key.service unit. On the shared box, sovrn’s metrics and relay and moods all declare keys, so names must not collide: each project’s secrets module prefixes the attribute (sovrn-metrics-password) and keeps the on-host file name through Colmena’s name option. Confirmed in Colmena 0.5.0’s source: the <attr>-key path and service units are named after the attribute, and the file after name. Services order on the unit name the module exposes (.unit), never a hard-coded string. The fleet’s servers.secrets already works this way (servers-<name>).

Within sovrn, metrics and relay share one key directory. Today their names don’t overlap (relay: stalwart-recovery-password, backup-domains-token, metrics-password; metrics: metrics-password-hash, pds-resend-api-key), but a name declared by two sovrn roles on one host must have identical settings, or the module system rejects it.

Toolkit used throughout

# What a host is running
ssh root@HOST readlink -f /run/current-system

# What the fleet would deploy
nix build ".#nixosConfigurations.\"HOST\".config.system.build.toplevel" -o result-HOST

# Compare (fetch the running closure first)
nix copy --from ssh-ng://root@HOST /nix/store/…-nixos-system-…
nix run nixpkgs#nvd -- diff /nix/store/…-nixos-system-… ./result-HOST

# Which units a deploy would touch, without touching them
colmena apply dry-activate --on HOST

# Deploy / deploy and reboot into it (build + boot + reboot)
colmena apply --on HOST
colmena apply --on HOST --reboot

# Roll back on the host
ssh root@HOST nixos-rebuild switch --rollback     # or pick the generation in GRUB

Wrap these as just diff-host HOST and just deploy HOST in the fleet’s Justfile. just deploy runs check-secrets HOST first, like sovrn and moods.


Phase 0: Survey and safety nets

Done:

Before each host’s first fleet deploy:

Phase 1: Fleet skeleton (no deploys) — done 2026-10-06

Built as below, with these differences from the original sketch:

Verified (evaluation only, nothing built or deployed):

The inventory is back to {}; hosts are added in 2.2 and 3.2.

Original sketch:

  1. Flake inputs:
    • nixpkgs = github:NixOS/nixpkgs/7a0f122f5090cf4c2ade2a13a0e229d4e19ba71f (sovrn’s and moods’ rev)
    • nixpkgs-stable = github:NixOS/nixpkgs/ac62194c3917d5f474c1a844b6fd6da2db95077d at first: exactly what infra.rtw.run runs, so 2.2 is a no-op. It moves to nixos-26.05 in 2.5.
    • colmena v0.5.0, disko, sovrn, moods
    • pgit and agenix locked to the revs in 8fd5074’s flake.lock (agenix is removed in 2.4)
  2. devShells.default: colmena, nixos-anywhere, nvd, jq, age, just. (secrets is ~/bin/secrets; the shell just needs it on PATH.)
  3. hive.nix: inventory → hive, modelled on moods’ nix/hosts.nix (nodeNixpkgs, nodeSpecialArgs, versionSuffix/revision, targetHost = name, targetUser = "root"). Expose colmenaHive and nixosConfigurations = colmenaHive.nodes.
  4. modules/: copy base.nix, vm.nix, hetzner.nix, netcup.nix, secrets.nix from moods; base.nix authorizes admins.pub for root. Not used by infra.rtw.run (legacy layout).
  5. nixpkgs-check.nix: assert inputs.sovrn.inputs.nixpkgs.rev == inputs.nixpkgs.rev and the same for moods, with a message that says to bump them together.
  6. just eval (Colmena eval of every node) passes with an empty inventory.

Pinning nixpkgs to the project rev keeps sovrn’s and moods’ promise that dev builds are host builds: the fleet applies the projects’ overlays to the same nixpkgs, so the store paths match.

Phase 2: infra.rtw.run → Colmena, secrets tool, current packages

In place, no reinstall. Each numbered deploy is separate. Order: adopt unchanged → keys → agenix out → upgrade → Caddy. The upgrade comes before the Caddy conversion so that conversion is written against the Caddy module version you’ll keep.

2.1 Put the secrets in the store

Done 2026-10-06: all four entries decrypt; pushover.env has both PUSHOVER_ variables, rclone.conf the [r2] remote, the hash is a crypt string, and the host key’s public half matches identities/host/infra-rtw-run.pub.

The plaintext sources (~/data) aren’t on the workstation, so read each value off the live host and pipe it straight into the store; nothing lands on local disk.

ssh [email protected] cat /etc/pushover.env \
  | secrets encrypt servers/shared/pushover.env
ssh [email protected] cat /root/.config/rclone/rclone.conf \
  | secrets encrypt servers/hosts/infra.rtw.run/rclone.conf
ssh [email protected] cat /etc/btburke-password \
  | secrets encrypt servers/hosts/infra.rtw.run/btburke-password-hash
ssh [email protected] cat /etc/ssh/ssh_host_ed25519_key \
  | secrets encrypt servers/hosts/infra.rtw.run/ssh_host_ed25519_key

If the R2 credentials in rclone.conf are the same as the store’s top-level rclone.conf.age, replace the host entry with a symlink to it instead (secrets follows links).

Check: secrets decrypt servers/shared/pushover.env | grep -c PUSHOVER_ → 2.

2.2 Adopt the host unchanged

  1. Add infra.rtw.run to hosts.json (layout: legacy, nixpkgs: stable, roles: [] for now, sshHostKey from identities/host/infra-rtw-run.pub).
  2. Its node imports exactly what 8fd5074’s flake.nix imports: agenix.nixosModules.default and ./hosts/infra-rtw-run, with the same specialArgs (flakeRoot, flakeHostName = "infra-rtw-run", pgit). Set system.nixos.versionSuffix/revision like moods’ hosts.nix, so the system label matches a nixosSystem build.
  3. deployment: targetHost = "infra.rtw.run", targetUser = "root", buildOnTarget = false. Builds moved to the server in 6b79956; only set buildOnTarget = true if the reason for that still applies.
  4. just known-hosts pins the key for both name and IP.

Done 2026-10-06: steps 1–4; just diff-host infra.rtw.run built the system locally and printed “identical” (r46cnzkj…); just known-hosts pinned the name (already that key) and the IP.

Verify it’s a no-op: just diff-host infra.rtw.run prints “identical” (already confirmed at evaluation in Phase 1). Then colmena apply dry-activate --on infra.rtw.run lists no units to restart.

  1. Root keys, separate commit. Add the hetzner key next to the scooter key in hosts/common/users.nix (or keyFiles = [ admins.pub ]), so it lands in /etc/ssh/authorized_keys.d/root. Kept out of the adoption commit so the no-op check above stays clean.

2.3 Deploy 1: switch deploy tooling

Done 2026-10-06:

Found during 2.3: leaked D-Bus daemons. Both activations printed reloading user units for root... Failed to open dbus connection, and [email protected] is failed with “Too many open files”. The post-receive hook’s pgit/pgit-index runs autolaunch a session dbus-daemon that never exits: 462 of them since 2026-06-23, each holding an inotify instance, which exhausted root’s fs.inotify.max_user_instances (128). Must be fixed before 2.4, whose Colmena key units run inotifywait as root. See 2.3a.

colmena apply --on infra.rtw.run

Nothing on the box changes. From now on the Makefile must not be used: delete the update, bootstrap, recover and reencrypt targets in the same commit (or the whole Makefile; the rest goes in 2.7).

Then deploy the root-key commit from 2.2 step 5 and check:

2.3a Stop the D-Bus leak (before 2.4)

Confirmed on the host: pgit-index with DBUS_SESSION_BUS_ADDRESS unset leaves one more orphaned dbus-daemon; with it set to disabled: it leaves none.

  1. Commit: the post-receive hook exports DBUS_SESSION_BUS_ADDRESS=disabled:. Only the hook (under /etc/soft-serve/hooks) and /etc change; nothing restarts.
  2. Deploy: just deploy infra.rtw.run.
  3. Clean up the orphans (root session buses only; the system bus runs as messagebus with --system): pkill -u 0 -f -- 'dbus-daemon --syslog --fork --print-pid 4 --print-address 6 --session', then systemctl reset-failed [email protected].
  4. Check (done 2026-10-06, generation 70):
    • [x] orphans gone: 463 → 0 (the only root dbus-daemon left is root’s real user bus, --address=systemd: under [email protected]); the system bus (messagebus) untouched
    • [x] [email protected] active, busctl --user works, no failed units; inotify instances on the box: 13 in total
    • [x] a repeat deploy activates without Failed to open dbus connection
    • [x] a push (the servers repo, 06:10) rebuilt the site index (13 repositories) and left no root dbus-daemon

Note for later: pkill -f <pattern> over ssh … '<command>' also matches the remote shell running the command (its argv contains the pattern) and kills it mid-way. Use pgrep/kill on PIDs, or a pattern trick like [d]bus-daemon.

2.4 Deploy 2: agenix → Colmena keys

Implemented 2026-10-06 (commit “agenix → Colmena keys”): hosts/common/secrets.nix imports modules/secrets.nix and declares pushover.env; hosts/infra-rtw-run/secrets.nix declares rclone.conf and btburke-password-hash; services order after <key>.unit. The agenix input, the legacy nixosConfigurations.infra-rtw-run output and agenix in the dev shell are gone (secrets/*.age stay in git until 2.7). Root also gets RCLONE_CONFIG in its environment, so interactive rclone and recover.sh find the config.

In hosts/infra-rtw-run/:

servers.secrets = {
  "pushover.env".scope = "shared";
  "rclone.conf".scope = "host";
  "btburke-password-hash".scope = "host";
};

All three are root-owned, so Colmena uploads them before activation; the password hash has to exist when the users step of activation reads it.

Was Becomes
age.secrets.*, agenix module, age.identityPaths removed
hashedPasswordFile = config.age.secrets.btburkePassword.path config.servers.secrets."btburke-password-hash".path
EnvironmentFile = "/etc/pushover.env" (soft-serve, backup, system-monitor) the pushover.env key path; wants/after its key unit
rclone reads /root/.config/rclone/rclone.conf backup.service: environment.RCLONE_CONFIG = the rclone.conf key path
recover.sh uses the default rclone config export RCLONE_CONFIG in it too, and fix its failure branch, which calls an undefined log

Users: servers doesn’t set users.mutableUsers = false, so NixOS keeps the existing password of an existing user, and swapping the hash file’s source doesn’t change btburke’s password. Leave mutableUsers alone on this host.

Verify before deploying: just diff-host shows agenix leaving and only the soft-serve, backup, system-monitor and user units changing. just check-secrets infra.rtw.run passes.

colmena apply --on infra.rtw.run --reboot

After it’s back:

Deployed 2026-10-06 with --reboot (generation 71, booted). A backup ran just before, on the agenix setup.

Rollback: the previous generation in GRUB, or colmena apply from the commit before. agenix’s ciphertext is still in git until 2.7.

2.5 Deploy 3: upgrade 25.05 → 26.05

soft-serve dry run (done 2026-10-06). 26.05 takes soft-serve from 0.8.5 to 0.11.6 (golang.org/x/crypto v0.49.0). The production data (/var/soft/data, 66 MB) was copied off the box and served locally by the 0.11.6 build from nixos-26.05, with the box’s config.yaml, local-only ports and the global hooks replaced by a test hook:

The copy (which included soft-serve’s private host key and the private repo) was deleted afterwards.

Rollback is still GRUB plus the step-2 backup, but since the database schema doesn’t change, 0.8.5 can read what 0.11.6 leaves behind.

This is goal 1’s package update: soft-serve, Caddy, the kernel and everything else move to 26.05. It skips 25.11.

  1. Read the 25.11 and 26.05 release notes for what this host uses through NixOS modules: openssh, users, networking (scripted networking.interfaces + dhcpcd), security.sudo, programs.neovim, GRUB. soft-serve and Caddy are hand-written units here, so their module changes don’t apply, but check soft-serve’s own changelog from 0.8.5 for config or database migrations.
  2. Fresh backup: systemctl start backup.
  3. Point nixpkgs-stable at github:NixOS/nixpkgs/nixos-26.05, nix flake update nixpkgs-stable. Also bump pgit if wanted (it’s built with the host’s Go, so it changes anyway).
  4. just diff-host infra.rtw.run: review the version changes (soft-serve, Caddy, kernel, systemd, openssh). Fix eval warnings and errors from renamed options.
  5. 
    colmena apply --on infra.rtw.run --reboot
    
  6. Check (deployed 2026-10-06 with --reboot, generation 72):
    • [x] nixos-version 26.05.20261006.b253099; soft-serve 0.11.6, Caddy 2.11.4, OpenSSH 10.5p1, kernel 6.18.55, systemd 260.4
    • [x] SSH as root (hetzner key and scooter key)
    • [x] IPv4: default route via 172.31.1.1, egress works. IPv6: see below
    • [x] git over SSH negotiates mlkem768x25519-sha256 (no post-quantum warning); git:// and HTTPS clone work
    • [x] https://kilimanjaro.io 200, www → 301; https://git.kilimanjaro.io/ is 404 straight from soft-serve (it has no page at /; repo paths return 200), HSTS header present
    • [x] a push returns at once (after 2.5b) and the background build regenerates the site (index rebuilt 07:12:41, 13 repos); no D-Bus orphans
    • [x] systemctl --failed is empty; backup and system-monitor run; no root dbus-daemon orphans

Caught before deploying: the default route. 26.05’s scripted networking only installs networking.defaultGateway on an interface it’s named for, or whose subnet contains the gateway. 172.31.1.1 is outside the /32, and defaultGateway had no interface, so the 26.05 build had no IPv4 default route at all; booted, the box would have been unreachable except through Hetzner’s console. (25.05’s network-setup.service added it unconditionally; 26.05 removed that unit.) Fixed by defaultGateway = { address = "172.31.1.1"; interface = "enp1s0"; }, checked in the generated network-addresses-enp1s0 script before deploying. The release note says “implementation details”; this is the kind of change only the generated units show.

Found: IPv6 has never worked on this box. There is no IPv6 default route (defaultGateway6 is unset; only fe80::1/128 is routed), so outbound IPv6 fails and the AAAA record for infra.rtw.run points at an address that can’t answer anyone off-link. kilimanjaro.io’s AAAA records are Cloudflare’s and git.kilimanjaro.io has none, so the services aren’t affected. Fix separately (2.5a).

2.5b Slow pushes after the upgrade (fixed 2026-10-06)

After 2.5 every push waited for the whole pgit site build (about a minute) even though the hook backgrounds it. Reproduced locally with a stand-in 15-second job: 0.2 s per push on soft-serve 0.8.5, 16.2 s on 0.11.6. The background job inherited fd 5, a pipe whose other end is held by the soft serve server process, which only ends the push at EOF on it. The fix in the hook: the background job closes every descriptor above 2 (and takes stdin from /dev/null) before doing anything. Locally: 0.2 s on both versions, every background build still completes.

2.5a IPv6 default route (done 2026-10-06, generation 74)

networking.defaultGateway6 = { address = "fe80::1"; interface = "enp1s0"; }; (Hetzner’s IPv6 gateway).

Not deployed with switch: changing network-addresses-enp1s0 restarts it, and its stop step deletes every address and route on the interface (IPv4 included) before the start step adds them back. Instead:

  1. Added the route by hand on the running box (ip -6 route replace default via fe80::1 dev enp1s0): IPv6 egress worked, and from infra.mymood.at (which has IPv6) inbound SSH, HTTP (308 to HTTPS) and git SSH on 23231 answered over IPv6.
  2. Deployed with --reboot.

After the reboot: both default routes present, curl -4 and curl -6 to example.com → 200, inbound IPv6 HTTP from infra.mymood.at → 308, no failed units. (Go’s SSH server on 23231 waits for the client’s version string before sending its own, so a bare banner probe looks like a dead port.)

Rollback: generation 69⁄70 in GRUB. soft-serve migrates its database forward on start, so rolling back across a soft-serve upgrade may need the R2 backup from step 2.

2.6 Deploy 4: Caddy onto services.caddy

Findings before implementing (2026-10-06):

Implemented in roles/forge.nix (infra.rtw.run gets roles: ["forge"]): the three sites verbatim; logFormat = null per site and level INFO globally to keep today’s logging (no access logs, renewals in the journal); a one-time caddy-storage-migrate unit copies the old storage into the new location before Caddy starts, so nothing is reissued. caddy adapt of the old and new Caddyfiles gives identical JSON apart from the source path and the explicit INFO level (the old default).

  1. Find where the current certificates live. The hand-written unit runs as root with ProtectSystem = "strict", writable /var/lib/caddy and no XDG_DATA_HOME: ssh [email protected] 'ls -R /var/lib/caddy | head; ls /root/.local/share/caddy'.
  2. Write roles/forge.nix with services.caddy.virtualHosts for git.kilimanjaro.io, www.kilimanjaro.io and kilimanjaro.io, each extraConfig taken from the current Caddyfile verbatim. Drop the requires = soft-serve.service coupling so the static site doesn’t go down with soft-serve.
  3. Remove systemd.services.caddy, environment.etc."caddy/Caddyfile", caddy from systemPackages, and the /var/lib/caddy tmpfiles rule.
  4. Certificates: either copy the existing ones into the module’s data directory (owned by caddy), or let Caddy issue fresh ones; three names are far under Let’s Encrypt’s limits, at the cost of a few seconds of TLS errors on the first requests.
colmena apply --on infra.rtw.run

Deployed 2026-10-06 (plain switch, generation 75):

Seen during activation, unrelated: efi.mount started. The disk carries an EFI system partition (sda15) from Hetzner’s cloud image, and systemd 260’s GPT auto-generator automounts it at /efi on first access. Harmless.

Left for 2.7: delete /root/.local/share/caddy and /root/.config/caddy, then drop the caddy-storage-migrate unit.

2.7 Move the rest into the forge role; remove the old tooling

  1. Move soft-serve, its config, the post-receive hook, the backup timer and the start notification from hosts/infra-rtw-run/services.nix into roles/forge.nix; set roles: ["forge"]. nvd diff should show no change.
  2. Fix the typo in config.yaml: ssh.public_url is ssh://git.kilimarjaro.io:23231 (should be kilimanjaro).
  3. Pull the site-generation loop out of the post-receive hook into a forge-rebuild-site command and have the hook call it. /var/www/code isn’t backed up, so after any restore this is how it comes back.
  4. Delete Makefile, scripts/, state/, secrets/, identities/, docs/agenix-rekey*.md, the agenix input and the devShell’s agenix/bash setup. Keep zone.db and recover.sh (move under roles/forge/). Rewrite README.md and AGENTS.md for the Colmena workflow.
  5. On the host: rm -r /etc/nixos (nixos-infect leftovers) and nix-channel --remove nixos.

Done 2026-10-06 (generation 76):

2.8 Optional later: rebuild on the standard layout

infra.rtw.run keeps hardware.nix (root mounted as /dev/sda1 by device name), the infect-era static networking and network-autodetect.nix. To get it onto vm.nix + hetzner.nix like every other host, without a reinstall in place:

A good moment for this is when the CI runner is added, if the box needs to be bigger anyway.

Phase 3: moods exports modules; fleet adopts infra.mymood.at

3.1 Refactor moods (in the moods repo)

Each step is verified by just build in moods producing an unchanged closure for infra.mymood.at.

  1. No more fleet/host specialArgs in service modules.
    • app.nix reads fleet // host.settings. Replace with an option moods.settings (attrs), defaulting to nix/fleet.nix via mkDefault.
    • secrets.nix uses host.name for host-scoped store paths. Use config.networking.fqdn (vm.nix sets hostName + domain). Two projects can’t both get a specialArg named fleet on one host: sovrn’s means mail domain/ACME/relay, moods’ means Jetstream/hostnames.
  2. Prefix key attributes (moods-anthropic-api-key, file name unchanged via name).
  3. Export nixosModules.services (moods.nix, secrets.nix) and nixosModules.app (app.nix, proxy.nix). Keep hosts.nix working on top of them until the fleet takes over.
  4. proxy.nix sets no global Caddy options; keep it that way (the fleet owns services.caddy.email).

3.2 Adopt the host

  1. Add infra.mymood.at to hosts.json (copy from moods’ inventory, roles: ["moods"]).
  2. just diff-host infra.mymood.at against the running system shows no differences. Differences here come from base/provider modules; fix the fleet copies until it’s clean.
  3. colmena apply --on infra.mymood.at (no-op).
  4. Remove deploy from moods’ Justfile, or make it print “deploy from ~/projects/servers: just deploy-project moods”.

From now on, shipping moods is:

cd ~/projects/servers
nix flake update moods          # picks up moods' last commit (jj: @-)
just deploy-project moods       # colmena apply --on @moods
jj commit -m "deploy moods <rev>"   # the lock bump is the deploy log

git+file inputs only see committed content.

3.3 Deploy: guard rails for sharing

Before adding anything else to the box:

Phase 3 record (done 2026-10-06)

Phases 4–11: sovrn onto the fleet, Ansible retired (revised 2026-10-06)

sovrn’s own epic (bug 3beadb2, “migrate provisioning from Ansible to NixOS + Colmena”) has its first seven children done (Hetzner provisioning, keys, packages, metrics pilot, Stalwart module, apply plan, orchestration). The earlier Phases 4–8 here covered only the relay out of what remains; this is the full path, mapped to the bugs. sovrn has no live hosts (mx99 and infra.sovrn.at were torn down), so nothing is migrated, only rebuilt.

Decisions behind it (2026-10-06):

Phase What Bugs
4 Key rotation; sovrn fleet integration; project Justfile stubs new fleet-integration child of 3beadb2, 3418287
5 Relay role b45a76f (d43eebe closes with it)
6 Telemetry edge a6edca3
7 Backups 0e32e5a
8 Metrics, relay and edge onto the shared box 6cae3a0 (partly)
9 mx99 rebuilt by the fleet; live acceptance; CI 63eb0d9, 6cae3a0, c3b0d04 (part)
10 Recovery, or an explicit decision to go without it for now 1badc33
11 Delete Ansible; cleanup in sovrn and moods c3b0d04

Not blocking Ansible’s retirement, but tracked: dfb37b1 (sovrnd /metrics; until it exists BackupStale and TlsExpiring can never fire, so do it before relying on alerts), edbcac6 (alerts on new forwarding channels), dc80d76 (Stalwart’s ASN/Geo download), af17761 (nixos-26.11, which now moves sovrn, moods and the shared box together).

Phase 4: Key rotation, sovrn fleet integration, Justfile stubs

  1. Rotate sovrn/shared/oidc-key.pem and sovrn/shared/oauth-attestation.key with sovrn’s own generator (go run ./cmd/gensecrets --dir <tmp>), into the store, the temp dir removed. Check nothing outside the store pins the old public halves (published JWKS, client metadata).
  2. Options instead of specialArgs. fleet (nix/fleet.nix) becomes sovrn.fleet with mkDefaults; host.name → config.networking.fqdn; host.pdsOriginSuffix → sovrn.cell.pdsOriginSuffix. Two projects can’t both get a specialArg named fleet.
  3. Hostnames as options. On the shared box the machine is infra.mymood.at, but metrics serves infra.sovrn.at and the relay is mxb.eu.sovrn.at: sovrn.metrics.hostname, sovrn.relay.hostname. Nothing in these roles may use the machine’s own name.
  4. Key names. Prefix key attributes sovrn- (files unchanged, as moods did); replace the hard-coded "pds-resend-api-key-key.service" and "metrics-password-hash-key.service" in roles/metrics.nix with a .unit attribute; metrics and relay must both fit on one host.
  5. Caddy email (roles/metrics.nix, cell-edge.nix): mkDefault, it’s one global value per host.
  6. Export nixosModules.services, .metrics, .cell (and .relay, .telemetryEdge as they land), each bringing the overlay.
  7. Roles from the inventory, not the hostname: the fleet’s hosts.json says sovrn-cell, sovrn-metrics, sovrn-relay.
  8. Host secrets generation moves to the fleet. The per-role generators (ssh_host_ed25519_key, Stalwart recovery password, ZDS secrets, …) stay defined by sovrn and are exposed from its flake (an app the fleet’s gen-host-secrets calls); new-host is merged into the fleet from sovrn’s and moods’ versions.
  9. Stubs. sovrn’s Justfile.nix host recipes (new-host, known-hosts, deploy, gen-host-secrets, check-secrets, ci-staging) and moods’ (new-host, known-hosts, deploy, build, eval-hosts, gen-host-secrets, check-secrets) call the fleet’s Justfile or print where to go.
  10. Verified by sovrn’s VM tests (cell, stalwart, metrics) and stalwart-plan; there is no running sovrn system to compare against.

Phase 4 record (done 2026-10-06).

Phase 5: Relay role (b45a76f)

nix/modules/roles/relay.nix, porting deployment/roles/stalwart_backup (install, bootstrap, routesync) and ADR-0009 D28 with its 2026-09 amendment:

Phase 5 record (done 2026-10-06).

Phase 6: Telemetry edge (a6edca3)

A sovrn.telemetry.edge module (exported for the fleet) on cells and on the shared box:

Phase 6 record (done 2026-10-07).

Phase 7: Backups (0e32e5a)

Litestream (stalwart.db, sovrn.db, oauth.db, ZDS DBs) and rclone (ZDS blobs every 15 min, cell config) to the cell’s R2 bucket, credentials as keys, health markers under /var/lib/sovrn/health with the ages sovrnd’s /healthz expects (15 m / 45 m). Done on a cell: markers fresh, R2 current.

Phase 7 record (done 2026-10-07).

Phase 8: Metrics, relay and edge on the shared box

Prerequisites on the box are in place (Phase 0: headroom, outbound 25 unblocked; Phase 3: guard rails).

  1. Metrics: DNS infra.sovrn.at → infra.mymood.at; add sovrn-metrics; secrets (metrics-password-hash, pds-resend-api-key) are in the store; dry-activate (moods must not restart), deploy. Check: https://infra.sovrn.at/healthz 200 and the Larm monitor pointed at it; vmui and logs UI behind basic auth; a test alert reaches [email protected] through Resend; the box still within its Phase 0 budget.
  2. Edge: on infra.mymood.at with the loopback endpoint; moods’ and sovrn’s units show up under cell="moods" / cell="sovrn".
  3. Relay: netcup PTR (IPv4 and IPv6) → mxb.eu.sovrn.at; DNS mxb.eu.sovrn.at → the box; per-host secrets via gen-host-secrets; add sovrn-relay, deploy. Check: from outside nc -v mxb.eu.sovrn.at 25 gets Stalwart’s banner (the inbound-25 test netcup’s policy might still block); STARTTLS presents a valid certificate; RCPT for any domain is rejected (closed, no cells yet); route-sync healthy.
  4. Leave backupMx = "" in sovrn’s fleet settings until Phase 9.

Phase 8 record (done 2026-10-07). Four deploys to infra.mymood.at, one change each; moods was never restarted.

Before Phase 9 (2026-10-07).

Phase 9: mx99 on the fleet; live acceptance

  1. new-host (fleet) builds mx99 fresh: roles: ["sovrn-cell"], exclusive.
  2. Live checks the epic has been waiting for:
    • the domain signup saga end to end on staging (63eb0d9)
    • mx99’s metrics and logs in VictoriaMetrics/Logs with labels the alert rules match (a6edca3); a staged failure fires the matching alert (6cae3a0)
    • backup markers fresh, R2 current (0e32e5a)
    • the relay pulls mx99’s domains; set backupMx = "mxb.eu.sovrn.at"; with mx99’s Stalwart stopped, mail queues on the relay and drains when it returns (b45a76f, d43eebe; docs/runbooks/relay-drain-verify.md)
  3. CI: sovrn’s ci-staging (tests, then deploy mx99) runs from this repo: bump the sovrn input, run sovrn’s checks, colmena apply --on mx99.eu.sovrn.at. Needs the CI remote URLs (open decision 1).

Phase 9 record (2026-10-07/08; CI step open).

Phase 10: Recovery (1badc33)

sovrn-restore.target and sovrn-freshen-backups as in the bug, rehearsed on a scratch cell. The Ansible playbooks are the only restore path today, so either this lands before Phase 11, or decide explicitly to delete Ansible without a restore path while there is no production data.

Phase 10 record (done 2026-10-08).

Phase 11: Decommission Ansible and clean up (c3b0d04)

Phase 11 record (done 2026-10-08).

Rollback summary

Step Rollback
2.3, 3.2 Nothing changed on the host
2.4, 2.6 Previous generation (GRUB or nixos-rebuild switch --rollback); agenix ciphertext in git until 2.7
2.5 Previous generation; soft-serve data from the step-2 R2 backup if its database migrated
3.3, 8 Remove the role from hosts.json and deploy; DNS for infra.sovrn.at / mxb.eu.sovrn.at can stay
4 (keys) None needed: nothing uses the keys yet
9 mx99 is staging: rebuild it

Risks and gotchas

Open decisions

  1. Where CI runs: none for now (2026-10-08); everything runs from the operator’s machine.
  2. buildOnTarget for infra.rtw.run.
  3. Whether and when to rebuild infra.rtw.run on the standard layout (2.8), e.g. together with adding the CI runner.

Decided: