T11: recovery-backup + recovery-bootstrap roles and playbooks
closedParent: bug 75966cc (cell architecture tracking). Split from T4 to keep replication infra shippable on its own; pairs with T10 (SSH posture) and the Justfile
recovery-backup/recovery-bootstraprecipes (already scaffolded — see Justfile; recipes fail clearly until these playbooks land).
Goal
The two recovery plays from docs/runbooks/recovery-primary-ip.md as real
Ansible: one read-mostly role against the sick box, one rebuild role against
its replacement. Same hostname, same recycled primary IPs, fresh host keys.
Scope
recovery-backuprole +playbooks/recover-backup.yml: SSH via.ssh/confighostname (strict host-key checking — this is the old known host; a changed key here is an incident, fail loudly). Stop services, force final Litestream sync +rclone sync --delete+stalwart-cli snapshotNDJSON, report backup ages. Failure-tolerant per task (failed_when: falsewhere sensible); never wipes.recovery-bootstraprole +playbooks/recover-bootstrap.yml: essentially themode=recoverpath renamed and scoped — install binaries,litestream restorex4 +PRAGMA integrity_checkeach, rclone sync-down, certs/config restored (never reissued), start in dependency order (Unbound -> Stalwart -> sovrnd -> ZDS -> Caddy -> Litestream/rclone), health gate. Host-key rotation is handled by the Justfile recipe before Ansible runs (ssh-keygen -R+ssh-keyscan); Ansible itself keeps strict checking on.- Manual middle step stays manual (operator powers off old, creates new box
re-linking the same
mxN.<region>.sovrn.at-{4,6}primary IPs, protection on, auto-delete off) — no provider API calls from Ansible.
Acceptance
- Staged drill on
mx99.eu.sovrn.atusing onlyjust recovery-backup/ manual move /just recovery-bootstrap: green healthz, relay drain clean, RTO/RPO posted on T9. T4 (replication roles) and T10 land first or with this.
2 Comments
Status: roles + playbooks implemented, drill + gates remaining
Implemented (commits
nsolvvxo,vrklvpol; full summary on T4 in comment2d5f579): - [x]recovery-backuprole +playbooks/recover-backup.yml— failure-tolerant freshen, never wipes; ZDS user units stopped via thepds-statusone-liner (not a system unit); WAL checkpoints on-target (no controller-side fileglob); best-effortstalwart-cliNDJSON (guarded — CLI likely absent on host). - [x]recovery-bootstraprole +playbooks/recover-bootstrap.yml— mirrorssite.ymlminus bootstrap-only destroy path; stop (install roles leave services running; restore is absent-only) →litestream-restore.sh(static x3 + dir loop) →integrity_checkgate → re-own per service user →rclone syncdown--delete→ start in dependency order. ZDS units left to sovrnd’s reconciler (no race). Strict host-key checking kept in Ansible; rotation stays in the Justfile recipe. - [x] Justfile stubs resolve (--checkclean on both playbooks).Remaining: - [ ] Staged drill on
mx99.eu.sovrn.at(acceptance: green healthz, relay drain clean, RTO/RPO on T9) — tracked as smoke-test item## 3on28c7d1f, blocked on the smoke server. - [ ] T10 (hetzner SSH posture) still missing — formally “lands first or with this”; soft dependency since key rotation already lives in the Justfile recipe. - [ ] Health gate (T6b957ddbopen) — no consolidated endpoint yet; drill uses interim Stalwart/healthz/ready+ ZDSdescribeServer+ manual JMAP login per the runbook. - [ ] Cert/config restore gap: scope says LE certs + Stalwart config are restored, never reissued — but the current backup set (SQLite via Litestream, blobs via rclone) covers neither. Decide: extend backup (e.g. rclone a config/certs dir) or consciously accept reissue on recovery (LE rate-limit risk on the runbook’s critical path). - [ ] Live-host confirmations from the drill: barelitestream snapshot -configscope, R2 generation sanity across the install-then-stop window, restored ownership.Item 4 closed: cert/config restore (
7d4261ed)cell-config-sync.sh(rclone role, runs first in the existing timer service): syncs/var/lib/caddy/(new varsovrn_caddy_storage) tocaddy-storage/with--delete, andcopytos/var/lib/stalwart/config.jsontostalwart-config/config.json. Guards skip absent paths so fresh boxes converge.recovery-backupfinal sync now runs config → blobs → Litestream snapshot.recovery-bootstraprestores both before units start: Caddy storage sync-down (best effort — a cert-less young cell degrades to on-demand reissue; chown gated onrc == 0, re-ownedcaddy:caddy0750/0600) and Stalwart configcopyto(fatal by design — without it Stalwart boots in bootstrap mode on a box where bootstrap is skipped).Remaining drill confirmations: cert files actually present under
/var/lib/caddypost-restore (no mass reissue in logs); Stalwart cert storage location confirmed — if its ACME certs live as files in the datadir outsidestalwart.db, extend the backup set to cover them.