1 files changed,
+36,
-1
+36,
-1
1@@ -1,6 +1,6 @@
2 # Fleet migration: sovrn, moods and servers → one Colmena fleet
3
4-Status: Phases 0–9 done except Phase 9's CI step (2026-10-08): infra.mymood.at runs moods, the metrics stack (infra.sovrn.at) and the backup relay (mxb.eu.sovrn.at); mx99.eu.sovrn.at is an arm64 staging cell on the fleet, accepted live (signup, mail, alerts, backups, relay drain). Next: Phase 10 (recovery).
5+Status: Phases 0–10 done except Phase 9's CI step (2026-10-08): infra.mymood.at runs moods, the metrics stack (infra.sovrn.at) and the backup relay (mxb.eu.sovrn.at); mx99.eu.sovrn.at is an arm64 staging cell, accepted live and rebuilt from its backups in a restore rehearsal. Next: Phase 11 (decommission Ansible).
6
7 ## Goals
8
9@@ -1417,6 +1417,41 @@ on a scratch cell. The Ansible playbooks are the only restore path today, so
10 either this lands before Phase 11, or decide explicitly to delete Ansible
11 without a restore path while there is no production data.
12
13+**Phase 10 record (done 2026-10-08).**
14+
15+- **Design (user):** restore is automatic and keyed off the backup guard,
16+ with no install flag. Cells keep their primary IPv4 and IPv6: Hetzner
17+ Primary IPs move to the replacement, so DNS and PTR never change.
18+- **sovrn `c6bfc5b`:** `sovrn-restore` (oneshot, before and required by
19+ Stalwart, sovrnd, Caddy and the backup units; waits for its R2 keys,
20+ retries until R2 answers). Armed or local databases present: nothing.
21+ Empty bucket: a new cell, armed. Otherwise: restore every database from
22+ Litestream (Stalwart, sovrnd's, `waitlist.db` on the landing cell, every
23+ ZDS database), `PRAGMA integrity_check` (any failure stops the unit, so no
24+ service starts), owners, ZDS blobs and Caddy storage copied back (never
25+ synced), arm. `sovrn-freshen-backups` stops the services and forces the
26+ final Litestream sync and rclone run on a cell about to be replaced.
27+ `checks.restore` (new) covers new cell, backup, freshen, empty replacement
28+ restore, continued backups, and never overwriting local data.
29+- **Rehearsal on mx99:** 5 messages in `[email protected]`;
30+ freshen; old CAX11 powered off; Primary IPs moved to a new CAX11; a Gmail
31+ message during the outage queued on the relay (route-sync kept the
32+ last-known domains); `just new-host` and `just deploy`; `sovrn-restore`
33+ restored all 5 databases before any service started; same accounts and
34+ domains; no new ACME orders (Stalwart and Caddy certificates came from the
35+ backup); ZDS back with the same DID and repo revision; `/healthz` ok; the
36+ relay delivered the outage message at its next retry; the phone shows all
37+ 6 messages.
38+- **Bug found and fixed (sovrn `9c9ee66`):** sovrnd didn't wait for the
39+ `zds-*` key units (Colmena uploads sovrn-owned keys after activation), so on
40+ first boot it skipped the ZDS env and left the restored instance down until
41+ restarted. Deployed.
42+- The relay's retry backoff grows (1, 5, 15 minutes), so mail sent during an
43+ outage can arrive up to ~15 minutes after the cell is back.
44+- Left for Phase 11: rewrite `docs/runbooks/recovery-primary-ip.md` and
45+ `pds-cell-recovery.md` for this flow; delete the old mx99 server once the
46+ user is satisfied.
47+
48 ### Phase 11: Decommission Ansible and clean up (`c3b0d04`)
49
50 - sovrn: delete `deployment/` (roles, playbooks, inventory, vault, scripts),