Recovery as systemd: sovrn-restore.target + freshen-backups
closedReplace playbooks/recover-backup.yml / recover-bootstrap.yml and roles/recovery-* with systemd units, so a rebuilt box restores itself before production services start.
Sketch:
- sovrn-restore.target with ordered oneshots, all Before= stalwart/sovrnd/caddy:
1. Litestream restore-if-absent (stalwart.db, sovrn.db, oauth.db, ZDS DBs)
2. PRAGMA integrity_check on all DBs; abort on anything but ok
3. Re-own files per service user
4. rclone sync of ZDS blobs down, then re-own
5. Caddy storage down (best effort)
- Gate: a flag file, kernel param, or NixOS specialisation (recovery), so normal boots never restore. Decide.
- After restore, Stalwart converges online, not in recovery mode, and never re-bootstraps.
- sovrn-freshen-backups oneshot (replaces recover-backup.yml): stop services, force final Litestream/rclone syncs, report backup ages. Triggered manually over SSH on the sick box.
- ZDS user units stay owned by sovrnd’s startup reconciler.
- Rewrite docs/runbooks/recovery-primary-ip.md and pds-cell-recovery.md.
Done when a rehearsal passes: build a scratch cell with data, freshen, destroy, nixos-anywhere a new box with the restore flag, then healthz is green with the data intact.
2 Comments
Built and rehearsed live (2026-10-08, fleet Phase 10). The done-condition passed on mx99.
Design (user): automatic, keyed off the backup guard, no flag. sovrn c6bfc5b: sovrn-restore runs before Stalwart, sovrnd, Caddy and the backup units: armed or local databases present -> nothing; empty bucket -> new cell, armed; otherwise restore every database (stalwart, sovrn, oauth, waitlist on the landing cell, every ZDS db), integrity_check, owners, copy ZDS blobs + Caddy storage back, arm. sovrn-freshen-backups for a cell being replaced. checks.restore (new) covers it. Cells keep their primary IPv4/IPv6 (Hetzner Primary IPs moved to the replacement), so no DNS/PTR changes.
Rehearsal: 5 messages in [email protected]; sovrn-freshen-backups at 05:27:40; old CAX11 powered off, Primary IPs moved to a new CAX11; one Gmail message during the outage (queued on the relay mxb.eu.sovrn.at; route-sync kept the last-known domains); just new-host + just deploy; sovrn-restore restored all 5 databases before any service started; same accounts and domains, no new ACME orders (Stalwart or Caddy: certificates from the backup), ZDS same DID and same repo rev 3mxdgsvxsxr27, /healthz ok; the relay delivered the outage message at its 05:52 retry; the phone shows all 6.
Bug found and fixed (sovrn 9c9ee66): sovrnd didn’t wait for the zds-* key units (Colmena uploads sovrn-owned keys after activation), so on first boot it skipped the ZDS env and left the restored instance down until restarted.
Notes: relay retry backoff grows (1, 5, 15 min), so outage mail can take ~15 min after the cell returns. Not done yet: rewrite docs/runbooks/recovery-primary-ip.md and pds-cell-recovery.md for this flow (Phase 11).
Closing (user, 2026-10-08): built (sovrn c6bfc5b, automatic restore keyed off the backup guard; sovrn-freshen-backups), rehearsed live on mx99 (rebuilt on a new box, everything restored, healthz green), and the runbooks rewritten (docs/runbooks/recovery-primary-ip.md, pds-cell-recovery.md).