plan: record Phase 10 (automatic restore, live rehearsal on mx99)
1 files changed,  +36, -1
M docs/fleet-migration-plan.md
+36, -1
 1@@ -1,6 +1,6 @@
 2 # Fleet migration: sovrn, moods and servers → one Colmena fleet
 3 
 4-Status: Phases 0–9 done except Phase 9's CI step (2026-10-08): infra.mymood.at runs moods, the metrics stack (infra.sovrn.at) and the backup relay (mxb.eu.sovrn.at); mx99.eu.sovrn.at is an arm64 staging cell on the fleet, accepted live (signup, mail, alerts, backups, relay drain). Next: Phase 10 (recovery).
 5+Status: Phases 0–10 done except Phase 9's CI step (2026-10-08): infra.mymood.at runs moods, the metrics stack (infra.sovrn.at) and the backup relay (mxb.eu.sovrn.at); mx99.eu.sovrn.at is an arm64 staging cell, accepted live and rebuilt from its backups in a restore rehearsal. Next: Phase 11 (decommission Ansible).
 6 
 7 ## Goals
 8 
 9@@ -1417,6 +1417,41 @@ on a scratch cell. The Ansible playbooks are the only restore path today, so
10 either this lands before Phase 11, or decide explicitly to delete Ansible
11 without a restore path while there is no production data.
12 
13+**Phase 10 record (done 2026-10-08).**
14+
15+- **Design (user):** restore is automatic and keyed off the backup guard,
16+  with no install flag. Cells keep their primary IPv4 and IPv6: Hetzner
17+  Primary IPs move to the replacement, so DNS and PTR never change.
18+- **sovrn `c6bfc5b`:** `sovrn-restore` (oneshot, before and required by
19+  Stalwart, sovrnd, Caddy and the backup units; waits for its R2 keys,
20+  retries until R2 answers). Armed or local databases present: nothing.
21+  Empty bucket: a new cell, armed. Otherwise: restore every database from
22+  Litestream (Stalwart, sovrnd's, `waitlist.db` on the landing cell, every
23+  ZDS database), `PRAGMA integrity_check` (any failure stops the unit, so no
24+  service starts), owners, ZDS blobs and Caddy storage copied back (never
25+  synced), arm. `sovrn-freshen-backups` stops the services and forces the
26+  final Litestream sync and rclone run on a cell about to be replaced.
27+  `checks.restore` (new) covers new cell, backup, freshen, empty replacement
28+  restore, continued backups, and never overwriting local data.
29+- **Rehearsal on mx99:** 5 messages in `[email protected]`;
30+  freshen; old CAX11 powered off; Primary IPs moved to a new CAX11; a Gmail
31+  message during the outage queued on the relay (route-sync kept the
32+  last-known domains); `just new-host` and `just deploy`; `sovrn-restore`
33+  restored all 5 databases before any service started; same accounts and
34+  domains; no new ACME orders (Stalwart and Caddy certificates came from the
35+  backup); ZDS back with the same DID and repo revision; `/healthz` ok; the
36+  relay delivered the outage message at its next retry; the phone shows all
37+  6 messages.
38+- **Bug found and fixed (sovrn `9c9ee66`):** sovrnd didn't wait for the
39+  `zds-*` key units (Colmena uploads sovrn-owned keys after activation), so on
40+  first boot it skipped the ZDS env and left the restored instance down until
41+  restarted. Deployed.
42+- The relay's retry backoff grows (1, 5, 15 minutes), so mail sent during an
43+  outage can arrive up to ~15 minutes after the cell is back.
44+- Left for Phase 11: rewrite `docs/runbooks/recovery-primary-ip.md` and
45+  `pds-cell-recovery.md` for this flow; delete the old mx99 server once the
46+  user is satisfied.
47+
48 ### Phase 11: Decommission Ansible and clean up (`c3b0d04`)
49 
50 - sovrn: delete `deployment/` (roles, playbooks, inventory, vault, scripts),