TRACKING: cell architecture (EU/US), floating IPs, SQLite+Litestream+R2, ZDS

closed
#75966cc opened by agent Sep 13

TRACKING: cell architecture — small cells, floating IPs, embedded DBs, ZDS

Supersedes: bug 6ca5959 (multi-MX/shared-FDB), bug 16df68d (imap./smtp. service names). Amends: bug 70790ed S5 (smoke scope), bug 9303770 (PDS pick: ZDS, E10 ops). Parent context: bug 3726817 (identity <-> mailbox mapping).

Decision (locked 2026-09-13)

One cell = 1 Hetzner box running Stalwart (SQLite) + sovrnd (SQLite: sovrn.db, oauth.db) + ZDS PDS (SQLite + disk blobs) + Unbound + Caddy. No FoundationDB, no Postgres, no coordinator. This supersedes ADR-0008 (active/passive FDB + VIP + leader lease).

Why cells: highest-probability failure mode is network issues; shared-store clusters (FDB clustering, PG streaming replication) convert every host failure into a network/consensus problem requiring expertise we don’t have in-house (no prod PG HA experience). Cells isolate blast radius to a handful of tenants / <100 mailboxes per box. ATProto identity is portable (PLC alsoKnownAs rotation, CAR export via PDSmoover/goat), so PDS-side exit and account moves stay well-tested and ZDS remains low-risk / replaceable.

Naming / regions: MX<number>.<region>.sovrn.at, regions EU/US for data locality (e.g. mx1.eu.sovrn.at, mx2.us.sovrn.at). Special-role boxes (outbound.eu.sovrn.at, metrics host) get their own roles later, once the SMTP-relay and monitoring decisions land. This issue covers cells only.

Failover: floating static IPs, never DNS. One floating IPv4 per cell (Hetzner, same network zone; guest-side /32 config in role; primary IP/IPv6 kept for provisioning + Ansible). Recovery: consolidated healthz unhealthy -> SSH if reachable -> stop services -> final backup flush -> new box same DC -> restore from R2 -> healthz green -> move floating IP via hcloud API -> power off old. MX/A records never change; the backup relay sees a momentary hiccup and drains its queue. Fencing before the move is mandatory (never move while old still serves).

Mail flow: inbound MX 10 -> cell floating IP; shared pri-20⁄30 relay(s) store-and-forward (Relay route pinned at cell, 5-7d queue). Outbound -> single warm relay; leaning 3rd-party, comail-first (ATProto-native, shared warmed pool) to avoid deliverability warmup. See relay-eval issue.

Backup/restore: one R2 bucket per cell named by hostname (e.g. mx1.eu.sovrn.at), manually created, R2 jurisdiction matching cell region. Layout inside the bucket: - litestream/<dbname>/... per SQLite file (Stalwart DB, sovrn.db, oauth.db, ZDS zds.sqlite3) — continuous, ~seconds RPO, PITR. - stalwart-blobs/... — Stalwart S3 blob store with keyPrefix: stalwart-blobs/ (content-hash keys, no tenant component). - zds-blobs/... — rclone sync target of the ZDS blob dir (interim), later the S3-PR object root blobs/{did}/{cid} + CDN origin. Each GC owns exactly its prefix (ZDS sweep, Stalwart purgeBlob); prefixes are never shared across cells or layers. Rclone uses --delete so S3 mirrors GC. Order: sync blobs first, snapshot DB second (orphan bytes self-heal via GC; missing bytes don’t). Plus stalwart-cli snapshot NDJSON to private repo (config layer) and optional periodic stalwart --export offline dumps.

ZDS blobs are all public (user ATProto activity, not sovrn; listable / downloadable, no auth — by design). Bot traffic is expected: long-term getBlob 302-redirects to public R2/CDN (cache-everything, immutable CIDs) so cells never pay read bandwidth. Mail blobs stay in a separate private bucket, JMAP-gated. Per-tenant blob erasure within a cell = account delete -> purgeBlob -> verify with list_active_blobs.py diff (Stalwart S3 keys are global content-hash objects; no per-tenant buckets in OSS).

Scale: manual. New customers -> newest cell with headroom; no tenant migration in v1. Per-cell ceilings (disk %, backup size/restore time, IMAP p99) documented as the “stand up another box” trigger.

Child issues

  • T1 dev PDS -> ZDS (devenv runner + smoke)
  • T2 Ansible ZDS role + Caddy/TLS + smoke deploy
  • T3 Stalwart -> SQLite, remove FDB (packaging / deploy / verify+docs)
  • T4 Litestream + rclone -> R2 (roles + secrets + restore drill)
  • T5 ZDS S3 blob + CDN redirect PR eval spike (Zig lib eval, then prototype)
  • T6 consolidated cell healthz in sovrnd
  • T7 metrics -> central VictoriaMetrics host (sovrn instrumentation + Stalwart/ZDS inventory + VM role/dashboards/alerts)
  • T8 SMTP relay eval (comail-first, Lettermint EU fallback)
  • T9 cell ADR + runbooks (provision / fenced floating-IP recovery / relay-drain verify)
  • T10 floating-IP + Hetzner role (guest /32, v6 bootstrap, known_hosts)

Sequence: T9 -> T3 -> T1 -> T2 -> T4 -> T6 -> T10 -> T7 -> T8 -> T5 (background; never blocks Sept — rclone covers launch).

Non-goals (v1)

Tenant migration tooling, active/active coordinator work, ZeroFS (dropped: S3-direct needs no POSIX bridge), per-tenant buckets, dynamic DNS failover.

2 Comments

agent 795f906 Sep 13

Revision (2026-09-13): primary-IP recycling, not floating IPs

Terminology correction locked across the plan: Hetzner Floating IPs cannot replace Primary IPs (same-type Primary required to carry one; no attach-at-creation; guest /32 always required). Cells standardize on Primary IPv4s as account resources (mxN.<region>.sovrn.at-{4,6}, protection on, auto-delete off), recycled across rebuilds. One IPv4 per cell; cloud-init configures primaries with zero guest config; MX/A and .ssh/config never change on recovery (only host keys rotate, handled by the recovery-bootstrap Justfile recipe). New T11 tracks the recovery role pair; T10 shrinks to SSH posture + known_hosts; “floating IP” language in T0/T9/ADR-0009/runbooks/deployment docs is being replaced with primary-recycle wording. Failover = rebuild + primary re-link (cold; old powered off first — matches the fencing requirement with no extra mechanism).

agent 7a5b966 Oct 8

Closing the tracker (2026-10-08): the cell architecture is built and running on NixOS and documented as it is in docs/deployment.md (that is now the source of truth; this issue’s text predates the move to recycled Primary IPs instead of floating IPs, - bucket names, no Unbound, and the end of Ansible). 13 of 16 descendants closed; still open, standalone: dfb37b1 (sovrnd /metrics) and the local/relay split verification (replaced by a fresh issue); 23e262e closed as moot.