T4: Litestream + rclone to R2 (roles, secrets, restore drill)
closedParent: bug 75966cc (cell architecture tracking). Buckets are manually created per cell, named by hostname (e.g.
mx1.eu.sovrn.at).
Goal
Continuous SQLite replication + blob sync to the cell’s R2 bucket, with a drilled restore path that the floating-IP recovery runbook depends on.
Scope (decompose as 4a/4b/4c/4d)
- 4a Litestream role: sidecar per SQLite file (Stalwart DB,
sovrn.db,oauth.db, ZDS DB) underlitestream/<dbname>/prefixes;litestream.ymltemplate; restore-if-missing on boot;snapshotsgate in deploy checks; PITR retention policy documented. - 4b rclone role: blob-dir sync with
--deletetozds-blobs/prefix (mirrors ZDS GC so S3 never resurrects); schedules; rclone remote config. Order invariant: sync blobs first, snapshot DB second. - 4c R2 secrets/config: vaulted endpoint/keys, per-cell bucket pattern,
retention/snapshot-interval policy in
sovrn.yml. - 4d restore drill: documented + staged rehearsal (kill cell, restore
latest,
PRAGMA integrity_check, healthz green). Relay queue covers the RPO gap.
Acceptance
litestream snapshotsnon-empty for every DB; PITR restore demonstrated in staging; runbook merged and drill passing.
2 Comments
SQLite + ZDS-blob Backup Implementation Plan (T4 + T11)
Goal: Continuous off-box backup for all SQLite files + the ZDS blob dir to the cell’s R2 bucket, with a drilled restore path that
just recovery-backup/bootstrapdepends on.Architecture: Litestream sidecars (3 static
pathentries for Stalwart/sovrn/oauth + 1dir+watch:trueentry for the ZDS per-domain DB dir, S3 backend to R2) for seconds-RPO + PITR; rclone systemd timer (/var/lib/zds/blobs/→zds-blobs/,--deleteso S3 mirrors GC); two T11 recovery roles (read-mostlyrecovery-backup, rebuildrecovery-bootstrap) wired to the existing Justfile stubs. Order invariant everywhere: blobs first, DB snapshot second (orphan bytes self-heal via GC; missing bytes don’t).Tech Stack: Litestream v0.5.x (static binary,
dir+pattern+watchdirectory watcher, S3-compatible R2 backend), rclone (S3 backend,sync --delete), Ansible roles + systemd service/timer, Cloudflare R2 (one bucket per cell), SQLitePRAGMA integrity_check / wal_checkpoint(TRUNCATE).Scope: T4 (bug 2559f24: 4a Litestream, 4b rclone, 4c R2 secrets/config, 4d restore drill) + T11 recovery role pair (bug c0e0ddb) in one slice, since recovery is undrillable without replication. Excludes: Stalwart S3
stalwart-blobs/live store wiring (stays docs-only), T5 S3-PR spike (23e262e), T6 healthz, T10 hetzner role, T8/T7.Parent context: bug 75966cc (cell architecture tracking); ADR-0009 D26 (bucket layout + ownership invariant); docs/deployment.md, docs/runbooks/provision-cell.md, docs/runbooks/recovery-primary-ip.md.
0. Manual operator steps for a new host (stays manual by design)
No provider API calls from Ansible. The operator does this per cell, once, in the Cloudflare dashboard + local vault ceremony:
mx1.eu.sovrn.at), jurisdiction matching cell region (EU cell → EU jurisdiction, US → US). No lifecycle rules; versioning off (Litestream owns generations under its prefix).https://<accountid>.r2.cloudflarestorage.com).mxN.<region>.sovrn.at-{4,6}(protection on, auto-delete off), set rDNS once, add.ssh/confighostname→IP entry. Same asdocs/runbooks/provision-cell.md §0.sovrn.yml. Vault it withansible-vault encrypt/ansible-vault editviadeployment/scripts/vault-password-client(same ceremony as relay/ZDS secrets).just build-stalwart+just build-zdsbefore bootstrap.litestream snapshotsnon-empty for every DB;rclone checkclean; healthz green. See Task 6.Why manual: bucket jurisdiction + token scoping are account-security actions with no safe service identity to automate under yet; a miscoped automation could cross cell prefixes and violate the ADR-0009 ownership invariant (each GC owns exactly its prefix; prefixes never shared across cells or layers).
File map (create / modify)
DB / path inventory (do not change these paths in this plan):
/var/lib/stalwart/stalwart.db(sovrn_stalwart_db)litestream/stalwart.db//var/lib/sovrn/sovrn.db(sovrn_datadir/sovrn.db)litestream/sovrn.db//var/lib/sovrn/oauth.dblitestream/oauth.db//var/lib/zds/dbs/<reversed>.db(e.g.com.example.alpha.db;internal/pdslifecycle/reverse.go,supervisor.go:29)litestream-zds/<reversed>.db/(auto-namespaced by dir watcher, see Task 2)/var/lib/zds/blobs/<slug>/(base/var/lib/zds/blobs;supervisor.go:30)zds-blobs/(rclone mirror,--delete)Task 1: R2/litestream/rclone config surface in
sovrn.ymlFiles: - Modify:
deployment/inventory/group_vars/all/sovrn.yml:43-46- Test:cd deployment && ansible-playbook playbooks/site.yml --check -l mx99.eu.sovrn.at(dry run resolves vars)host_vars/<cell>/vault.ymlonly):Run:
cd deployment && ansible-playbook playbooks/site.yml --check -l mx99.eu.sovrn.at 2>&1 | head -n 40Expected: noundefined variableforsovrn_r2_*/sovrn_litestream_*.Task 2:
litestreamrole — continuous SQLite replication (static + dir-watch)Files: - Create:
deployment/roles/litestream/tasks/main.yml- Create:deployment/roles/litestream/handlers/main.yml- Create:deployment/roles/litestream/templates/litestream.yml.j2- Create:deployment/roles/litestream/templates/litestream-restore.sh.j2- Modify:deployment/playbooks/site.yml(add role +litestreamtag)Design (locked): - Static binary install, pinned
sovrn_litestream_version(v0.5.x withwatchsupport; fill checksum at implementation), checksum-verified; no apt repo. - 3 staticpathentries (stalwart/sovrn/oauth) + 1dir+watch:trueentry for ZDS:dir: /var/lib/zds/dbs,pattern: "*.db",watch: true,replica url: s3://<bucket>/litestream-zds. The watcher (fsnotify, fine on Debian) validates SQLite headers and starts replication within seconds of the PDS provisioner creating a new<reversed>.db— no Ansible re-render or new-domain hook needed. Empty dir on startup is allowed withwatch: true.recursive: false(flat dir). - Replica paths for dir entries namespace automatically by relative path (com.example.alpha.db→litestream-zds/com.example.alpha.db/ltx/...); document this layout since restore scripts must match it. -restore-if-missing on boot:ExecStartPre=/usr/local/sbin/litestream-restore.shrestores only when the local file is absent (never clobbers live data; the T11mode=recoverpath does the authoritative restore).litestream.yml.j2litestream-restore.sh.j2(static restores + dir loop)site.yml(aftersovrnd/zds, taglitestream; zds role must create/var/lib/zds/dbsbefore litestream starts).Run:
just update mx99.eu.sovrn.at --tags litestreamThen:ssh mx99.eu.sovrn.at 'sudo litestream snapshots -config /etc/litestream.yml 2>&1 | head -n 20'Expected: one non-empty snapshot line per static DB plus one per<reversed>.dbunderlitestream-zds/.litestream snapshotsgains the new<reversed>.dbwithin ~1 min with no Ansible run.Task 3:
rclonerole — ZDS blob dir →zds-blobs/Files: - Create:
deployment/roles/rclone/tasks/main.yml- Create:deployment/roles/rclone/templates/rclone.conf.j2- Create:deployment/roles/rclone/templates/zds-blobs-sync.sh.j2- Create:deployment/roles/rclone/templates/zds-blobs-sync.service.j2- Create:deployment/roles/rclone/templates/zds-blobs-sync.timer.j2Invariants:
--delete(S3 mirrors ZDS sweep GC, never resurrects); timer default every 15 min viasovrn_rclone_schedule; on-demand fromrecovery-backup. Runbook order always blobs before DB snapshots.rclone.conf.j2+zds-blobs-sync.sh.j2+ unitsRun:
just update mx99.eu.sovrn.at --tags rclone && ssh mx99.eu.sovrn.at 'sudo /usr/local/sbin/zds-blobs-sync.sh && rclone lsd r2:mx99.eu.sovrn.at/zds-blobs/ --config /etc/rclone.conf'Expected: exit 0,zds-blobs/lists per-slug dirs.Task 4: T11
recovery-backup+recovery-bootstraproles and playbooksFiles: - Create:
deployment/roles/recovery-backup/tasks/main.yml- Create:deployment/roles/recovery-bootstrap/tasks/main.yml- Create:deployment/playbooks/recover-backup.yml- Create:deployment/playbooks/recover-bootstrap.ymlThese satisfy the existing Justfile stubs (
Justfile:152-177currently fail with “not yet implemented”). Strict host-key checking stays on in Ansible; key rotation lives only in therecovery-bootstrapJustfile recipe (already scaffolded). The manual provider middle step (power off old, create new box re-linking the samemxN.<region>.sovrn.at-{4,6}primaries) stays manual — no provider API from Ansible.recovery-backup(failure-tolerant, never wipes)recovery-bootstrap(authoritative restore, aborts onintegrity_check!= ok)Run:
just recovery-backup mx99.eu.sovrn.at -- --checkandjust recovery-bootstrap mx99.eu.sovrn.at -- --checkExpected: playbook found (no “not yet implemented” error); check-mode green.Task 5: Staged restore drill + runbook acceptance (mx99)
Files: - Modify:
docs/runbooks/provision-cell.md(clear TODO(T4/T11) block, link new roles) - Test: live drill onmx99.eu.sovrn.atjust recovery-backup mx99.eu.sovrn.at, manual provider move (power off old, new box re-linking primaries),just recovery-bootstrap mx99.eu.sovrn.at, then:Expected: every
integrity_check→ok; snapshots non-empty per DB (static + each<reversed>.dbunderlitestream-zds/);rclone checkexit 0; healthz green; relay-drain verify perrelay-drain-verify.md; post RTO/RPO on T9.TODO(T4)/TODO(T11)block inprovision-cell.md:76-85, replace with role names + manual-bucket reminder from §0.Self-review
sovrn_r2_*,sovrn_litestream_*,sovrn_zds_*) match across templates and tasks; R2 prefixes (litestream/,litestream-zds/,zds-blobs/) match restore scripts and runbooks.Implementation complete (Tasks 1-5, untested on live host)
Stack on top of
tzkupqok: -pukoovwpfeat(backup): R2/litestream/rclone config surface (sovrn.ymldefaults +host_vars/mx99.eu.sovrn.at/vars.yml; litestream 0.5.13 checksum resolved from upstream) -lopsvlnpfeat(backup): litestream role — 3 static DBs +dir/watch:truefor/var/lib/zds/dbs(new ZDS DBs replicate automatically, no re-render hook), restore-if-missing helper, wired intosite.yml-vonzoqoxfeat(backup): rclone role —zds-blobs/sync--deleteon systemd timer, wired intosite.yml-nsolvvxofeat(backup): T11recovery-backup(failure-tolerant freshen, never wipes) +recovery-bootstrap(authoritative restore,integrity_checkgate, re-own per service user) +recover-backup.yml/recover-bootstrap.yml; satisfies the Justfile stubs -ukywsxssdocs(backup): cleared stale TODO(T4)/TODO(T11) in provision-cell.md; deployment.md layout marked accurate -vrklvpolchore(backup): stub guards now say “missing” instead of “not yet implemented”Verify:
ansible-playbook --syntax-checkclean on all 3 playbooks;go build ./...OK.Manual / untested (operator): - R2 bucket + bucket-scoped token creation per cell stays manual (keys into
host_vars/<cell>/vault.yml). - Live restore drill onmx99.eu.sovrn.atnot run — do it before pointing MX at any cell. - Known caveats for the drill:litestream snapshot(no path arg) assumed v0.5.x-wide;stalwart-cliNDJSON step is a guarded no-op until the CLI story settles; confirm R2 generations stay sane on first recovery (install-then-stop window).