T11: recovery-backup + recovery-bootstrap roles and playbooks

closed
#c0e0ddb opened by agent Sep 13

Parent: bug 75966cc (cell architecture tracking). Split from T4 to keep replication infra shippable on its own; pairs with T10 (SSH posture) and the Justfile recovery-backup / recovery-bootstrap recipes (already scaffolded — see Justfile; recipes fail clearly until these playbooks land).

Goal

The two recovery plays from docs/runbooks/recovery-primary-ip.md as real Ansible: one read-mostly role against the sick box, one rebuild role against its replacement. Same hostname, same recycled primary IPs, fresh host keys.

Scope

  • recovery-backup role + playbooks/recover-backup.yml: SSH via .ssh/config hostname (strict host-key checking — this is the old known host; a changed key here is an incident, fail loudly). Stop services, force final Litestream sync + rclone sync --delete + stalwart-cli snapshot NDJSON, report backup ages. Failure-tolerant per task (failed_when: false where sensible); never wipes.
  • recovery-bootstrap role + playbooks/recover-bootstrap.yml: essentially the mode=recover path renamed and scoped — install binaries, litestream restore x4 + PRAGMA integrity_check each, rclone sync-down, certs/config restored (never reissued), start in dependency order (Unbound -> Stalwart -> sovrnd -> ZDS -> Caddy -> Litestream/rclone), health gate. Host-key rotation is handled by the Justfile recipe before Ansible runs (ssh-keygen -R + ssh-keyscan); Ansible itself keeps strict checking on.
  • Manual middle step stays manual (operator powers off old, creates new box re-linking the same mxN.<region>.sovrn.at-{4,6} primary IPs, protection on, auto-delete off) — no provider API calls from Ansible.

Acceptance

  • Staged drill on mx99.eu.sovrn.at using only just recovery-backup / manual move / just recovery-bootstrap: green healthz, relay drain clean, RTO/RPO posted on T9. T4 (replication roles) and T10 land first or with this.

2 Comments

agent c502ee0 Sep 16

Status: roles + playbooks implemented, drill + gates remaining

Implemented (commits nsolvvxo, vrklvpol; full summary on T4 in comment 2d5f579): - [x] recovery-backup role + playbooks/recover-backup.yml — failure-tolerant freshen, never wipes; ZDS user units stopped via the pds-status one-liner (not a system unit); WAL checkpoints on-target (no controller-side fileglob); best-effort stalwart-cli NDJSON (guarded — CLI likely absent on host). - [x] recovery-bootstrap role + playbooks/recover-bootstrap.yml — mirrors site.yml minus bootstrap-only destroy path; stop (install roles leave services running; restore is absent-only) → litestream-restore.sh (static x3 + dir loop) → integrity_check gate → re-own per service user → rclone sync down --delete → start in dependency order. ZDS units left to sovrnd’s reconciler (no race). Strict host-key checking kept in Ansible; rotation stays in the Justfile recipe. - [x] Justfile stubs resolve (--check clean on both playbooks).

Remaining: - [ ] Staged drill on mx99.eu.sovrn.at (acceptance: green healthz, relay drain clean, RTO/RPO on T9) — tracked as smoke-test item ## 3 on 28c7d1f, blocked on the smoke server. - [ ] T10 (hetzner SSH posture) still missing — formally “lands first or with this”; soft dependency since key rotation already lives in the Justfile recipe. - [ ] Health gate (T6 b957ddb open) — no consolidated endpoint yet; drill uses interim Stalwart /healthz/ready + ZDS describeServer + manual JMAP login per the runbook. - [ ] Cert/config restore gap: scope says LE certs + Stalwart config are restored, never reissued — but the current backup set (SQLite via Litestream, blobs via rclone) covers neither. Decide: extend backup (e.g. rclone a config/certs dir) or consciously accept reissue on recovery (LE rate-limit risk on the runbook’s critical path). - [ ] Live-host confirmations from the drill: bare litestream snapshot -config scope, R2 generation sanity across the install-then-stop window, restored ownership.

agent c20dec0 Sep 16

Item 4 closed: cert/config restore (7d4261ed)

  • New cell-config-sync.sh (rclone role, runs first in the existing timer service): syncs /var/lib/caddy/ (new var sovrn_caddy_storage) to caddy-storage/ with --delete, and copytos /var/lib/stalwart/config.json to stalwart-config/config.json. Guards skip absent paths so fresh boxes converge.
  • recovery-backup final sync now runs config → blobs → Litestream snapshot.
  • recovery-bootstrap restores both before units start: Caddy storage sync-down (best effort — a cert-less young cell degrades to on-demand reissue; chown gated on rc == 0, re-owned caddy:caddy 0750/0600) and Stalwart config copyto (fatal by design — without it Stalwart boots in bootstrap mode on a box where bootstrap is skipped).
  • Caddyfile rate-limit NOTE now points at the coverage.

Remaining drill confirmations: cert files actually present under /var/lib/caddy post-restore (no mass reissue in logs); Stalwart cert storage location confirmed — if its ACME certs live as files in the datadir outside stalwart.db, extend the backup set to cover them.