TRACKING: cell architecture (EU/US), floating IPs, SQLite+Litestream+R2, ZDS
closedTRACKING: cell architecture — small cells, floating IPs, embedded DBs, ZDS
Supersedes: bug 6ca5959 (multi-MX/shared-FDB), bug 16df68d (imap./smtp. service names). Amends: bug 70790ed S5 (smoke scope), bug 9303770 (PDS pick: ZDS, E10 ops). Parent context: bug 3726817 (identity <-> mailbox mapping).
Decision (locked 2026-09-13)
One cell = 1 Hetzner box running Stalwart (SQLite) + sovrnd (SQLite:
sovrn.db, oauth.db) + ZDS PDS (SQLite + disk blobs) + Unbound + Caddy.
No FoundationDB, no Postgres, no coordinator. This supersedes ADR-0008
(active/passive FDB + VIP + leader lease).
Why cells: highest-probability failure mode is network issues; shared-store
clusters (FDB clustering, PG streaming replication) convert every host failure
into a network/consensus problem requiring expertise we don’t have in-house
(no prod PG HA experience). Cells isolate blast radius to a handful of
tenants / <100 mailboxes per box. ATProto identity is portable (PLC
alsoKnownAs rotation, CAR export via PDSmoover/goat), so PDS-side exit and
account moves stay well-tested and ZDS remains low-risk / replaceable.
Naming / regions: MX<number>.<region>.sovrn.at, regions EU/US for data
locality (e.g. mx1.eu.sovrn.at, mx2.us.sovrn.at). Special-role boxes
(outbound.eu.sovrn.at, metrics host) get their own roles later, once the
SMTP-relay and monitoring decisions land. This issue covers cells only.
Failover: floating static IPs, never DNS. One floating IPv4 per cell (Hetzner, same network zone; guest-side /32 config in role; primary IP/IPv6 kept for provisioning + Ansible). Recovery: consolidated healthz unhealthy -> SSH if reachable -> stop services -> final backup flush -> new box same DC -> restore from R2 -> healthz green -> move floating IP via hcloud API -> power off old. MX/A records never change; the backup relay sees a momentary hiccup and drains its queue. Fencing before the move is mandatory (never move while old still serves).
Mail flow: inbound MX 10 -> cell floating IP; shared pri-20⁄30 relay(s) store-and-forward (Relay route pinned at cell, 5-7d queue). Outbound -> single warm relay; leaning 3rd-party, comail-first (ATProto-native, shared warmed pool) to avoid deliverability warmup. See relay-eval issue.
Backup/restore: one R2 bucket per cell named by hostname
(e.g. mx1.eu.sovrn.at), manually created, R2 jurisdiction matching cell
region. Layout inside the bucket:
- litestream/<dbname>/... per SQLite file (Stalwart DB, sovrn.db,
oauth.db, ZDS zds.sqlite3) — continuous, ~seconds RPO, PITR.
- stalwart-blobs/... — Stalwart S3 blob store with
keyPrefix: stalwart-blobs/ (content-hash keys, no tenant component).
- zds-blobs/... — rclone sync target of the ZDS blob dir (interim), later
the S3-PR object root blobs/{did}/{cid} + CDN origin.
Each GC owns exactly its prefix (ZDS sweep, Stalwart purgeBlob); prefixes
are never shared across cells or layers. Rclone uses --delete so S3 mirrors
GC. Order: sync blobs first, snapshot DB second (orphan bytes self-heal via
GC; missing bytes don’t). Plus stalwart-cli snapshot NDJSON to private repo
(config layer) and optional periodic stalwart --export offline dumps.
ZDS blobs are all public (user ATProto activity, not sovrn; listable /
downloadable, no auth — by design). Bot traffic is expected: long-term
getBlob 302-redirects to public R2/CDN (cache-everything, immutable CIDs)
so cells never pay read bandwidth. Mail blobs stay in a separate private
bucket, JMAP-gated. Per-tenant blob erasure within a cell = account delete
-> purgeBlob -> verify with list_active_blobs.py diff (Stalwart S3 keys
are global content-hash objects; no per-tenant buckets in OSS).
Scale: manual. New customers -> newest cell with headroom; no tenant migration in v1. Per-cell ceilings (disk %, backup size/restore time, IMAP p99) documented as the “stand up another box” trigger.
Child issues
- T1 dev PDS -> ZDS (devenv runner + smoke)
- T2 Ansible ZDS role + Caddy/TLS + smoke deploy
- T3 Stalwart -> SQLite, remove FDB (packaging / deploy / verify+docs)
- T4 Litestream + rclone -> R2 (roles + secrets + restore drill)
- T5 ZDS S3 blob + CDN redirect PR eval spike (Zig lib eval, then prototype)
- T6 consolidated cell healthz in sovrnd
- T7 metrics -> central VictoriaMetrics host (sovrn instrumentation + Stalwart/ZDS inventory + VM role/dashboards/alerts)
- T8 SMTP relay eval (comail-first, Lettermint EU fallback)
- T9 cell ADR + runbooks (provision / fenced floating-IP recovery / relay-drain verify)
- T10 floating-IP + Hetzner role (guest /32, v6 bootstrap, known_hosts)
Sequence: T9 -> T3 -> T1 -> T2 -> T4 -> T6 -> T10 -> T7 -> T8 -> T5 (background; never blocks Sept — rclone covers launch).
Non-goals (v1)
Tenant migration tooling, active/active coordinator work, ZeroFS (dropped: S3-direct needs no POSIX bridge), per-tenant buckets, dynamic DNS failover.
2 Comments
Revision (2026-09-13): primary-IP recycling, not floating IPs
Terminology correction locked across the plan: Hetzner Floating IPs cannot replace Primary IPs (same-type Primary required to carry one; no attach-at-creation; guest /32 always required). Cells standardize on Primary IPv4s as account resources (
mxN.<region>.sovrn.at-{4,6}, protection on, auto-delete off), recycled across rebuilds. One IPv4 per cell; cloud-init configures primaries with zero guest config; MX/A and.ssh/confignever change on recovery (only host keys rotate, handled by therecovery-bootstrapJustfile recipe). New T11 tracks the recovery role pair; T10 shrinks to SSH posture + known_hosts; “floating IP” language in T0/T9/ADR-0009/runbooks/deployment docs is being replaced with primary-recycle wording. Failover = rebuild + primary re-link (cold; old powered off first — matches the fencing requirement with no extra mechanism).Closing the tracker (2026-10-08): the cell architecture is built and running on NixOS and documented as it is in docs/deployment.md (that is now the source of truth; this issue’s text predates the move to recycled Primary IPs instead of floating IPs,-| bucket names, no Unbound, and the end of Ansible). 13 of 16 descendants closed; still open, standalone: dfb37b1 (sovrnd /metrics) and the local/relay split verification (replaced by a fresh issue); 23e262e closed as moot. |