T6: consolidated cell healthz in sovrnd
closedParent: bug 75966cc (cell architecture tracking). Gates floating-IP moves; keep metrics out (see metrics issue).
Goal
One endpoint expressing whole-cell health, suitable for paging and as the pre-move gate in the floating-IP recovery runbook.
Scope
GET /healthz/ready(+/live): aggregate with per-check timeouts and cached results — Stalwart ready + authed JMAP probe, ZDSdescribeServer, Unbound test resolutions (cell MX + external name), SQLite writability, disk + Litestream/rclone backup freshness.- JSON per-component body for humans, binary status code for automation.
- Unit + integration tests; thresholds/check list documented.
Acceptance
- Single endpoint reflects cell state truthfully under staged failures (each dependency down => non-OK with identifying body); docs merged.
3 Comments
Consolidated
GET /healthzImplementation Plan (bug b957ddb)Goal: Add unauthenticated
GET /healthzto sovrnd returning 200 + per-check JSON when healthy, 503 + same JSON shape when any dependency fails.Architecture: New
healthz.go(packagesovrn, next tocaddy_ask.go) with an injectable checker set run concurrently under per-check timeouts; mounted on the unauthenticatedtopmux inrouter.go(not the authed XRPC mux). No caching in v1 — concurrency + tight timeouts keep it under Larm’s 10s probe timeout; Larm’s multi-location majority voting means the handler must be concurrency-safe.Tech Stack: Go stdlib
net/http+http.ServeMux, existingstalwart.Client(JMAP),store.Store(SQLite),domain.ResolverDNS, viper config,httptesttests.Decisions locked: path
GET /healthzonly (not/healthz/ready); 503 on any failure; unauthenticated public; endpoint-only scope (Larm monitor configured manually — recommended settings documented in Task 5).Larm finding (https://docs.larm.dev/monitors): HTTP monitor,
Method: GET,URL: https://<cell>/healthz,Expected status codes: [200]— status code is the sole error signal (no keyword needed; body shape identical on success/failure).Timeout: 10s,Confirm down: 1 min,Confirm up: 3 mindefaults; no automated recovery — alerts only, operator runsdocs/runbooks/recovery-primary-ip.md.Task 1: Health response types + handler skeleton
Files: - Create:
healthz.go- Test:healthz_test.goTestHealthzAllGreen— single ok checker → 200 +status: okJSON).go test -run TestHealthzAllGreen ., expect FAIL).checkResult,healthResponse,checkFunc,healthHandler.ServeHTTPwith concurrent per-check timeouts, 200⁄503 aggregation,application/json).go test -run TestHealthzAllGreen ., expect PASS).Task 2: Failure semantics — any fail → 503, same body
Files: - Modify:
healthz_test.go(staged-failure matrix over stalwart/zds/dns/sqlite/disk/backup)TestHealthzAnyFailureIs503(each single failing check → 503 +status: degraded+ identifyingerror).Task 3: Real checkers — Stalwart, ZDS, DNS, SQLite, disk, backup
Files: - Modify:
healthz.go(constructors),healthz_test.go(fake-server tests),config.go(+HealthConfig, defaults, bindEnv),config_test.go,deployment/roles/sovrnd/templates/sovrn.toml.j2([health]block)GET {url}/healthz/ready+ authed JMAPQuery(Tenant); zds =GET {zdsurl}/xrpc/com.atproto.server.describeServer; dns = resolve cell MX + external name via cell Unbound; sqlite = writability probe (SELECT + transactional write/rollback, no residue); disk = statfs used% vs threshold (default 85); backup = Litestream/rclone marker mtime vs max age (default 10m).HealthConfig{PerCheckTimeout 2s, ZDSURL, CellMX, ExternalName example.com, DiskPath=DataDir, DiskMaxUsedPct 85, BackupMaxAge 10m, LitestreamState}.Task 4: Wire
GET /healthzinto the router (unauthenticated, public)Files: - Modify:
router.go(top.Handle("GET /healthz", logging.Middleware(newHealthHandler(cfg, st, sw), nil))before/catch-all), router test (open endpoint, JSON content-type, 200-or-503),sovrn.toml.j2+ inventory vars.Task 5: Docs — check list, thresholds, Larm monitor recipe
Files: - Modify:
docs/deployment.md,docs/runbooks/provision-cell.md,docs/runbooks/recovery-primary-ip.md,sovrn.toml.j2comments.Document each check, failure meaning, defaults, and the Larm recipe (URL, GET, expected
[200], confirm down 1m/up 3m, alert channels; no recovery webhook). Final:go test ./...+go vet ./...green.Implementation complete: consolidated
GET /healthzImplemented per the plan above (comment bf93567). All in the working copy, uncommitted (the tree already held unrelated backup/Caddy work — left untouched).
What was built
healthz.go(new): unauthenticatedGET /healthz, 200 +{"status":"ok",...}when healthy, 503 +{"status":"degraded",...}with the same per-check shape otherwise. Six checks, run concurrently underhealth.perchecktimeout(2s):stalwart:GET {stalwart.url}/healthz/ready+ authed JMAPx:Tenant/queryprobe (either leg names itself on failure)zds: every active PDS instance from the registry,describeServeron its loopback port with origin verification (supervisor’s healthGate semantics). Supervision OFF skips ok (caddy-ask rationale); no actives skips ok. No static URL — ZDS is per-domain instances, so the plannedhealth.zdsurlwas dropped.dns: resolves configured cell MX + external names via system resolver (Unbound on cells); names failuressqlite:Store.Ping— SELECT plus rolled-back INSERT (no residue)disk: statfs used% vs threshold (default 85)backup: marker-file mtime vs max age (default 10m); empty marker skips “unconfigured” ok until the sync scripts touch it (follow-up, noted in provision-cell.md)router.go: mounted on the unauthenticated top mux before the/catch-all; Caddy needs no change (only/internal/*is blocked at the edge).config.go+sovrn.toml.j2+sovrn.yml: new[health]section (perchecktimeout,cellmxdefaulting to mail hostname,externalnamedefaultexample.com,diskmaxusedpct,backupmaxage,backupmarkerfile).internal/store:Pingadded to theStoreinterface (sole impl is sqlite).deployment.md(check list, thresholds, Larm recipe), both runbooks now gate onGET /healthzinstead of the interim manual checks.Larm setup (manual, endpoint-only scope)
HTTP monitor:
GET https://<cell>/healthz, expected status[200], timeout 10s, confirm down 1m / up 3m. Status code is the sole signal — no keyword matching (passing checks also contain “ok”). Alerts only; recovery stays manual per the runbook.Verification
NewApp(503 naming dead Stalwart).go vet ./...clean;SOVRN_INTEGRATION=0 go test ./...all packages pass; root package also passes under-race.Leaving open for your review — close when satisfied.
Backup markers wired: litestream poller + rclone touch
The
backupleg is now two live checks instead of “unconfigured”.How freshness is proven
backup-litestream): continuousreplicatehas no periodic hook, so a new 5-min timer (litestream-freshness.timer,sovrn_litestream_freshness_schedule: "*:0/5") runslitestream-freshness.sh, which executeslitestream sync -wait -timeout 30 <db>per database (stalwart.db, sovrn.db, oauth.db, ZDS*.db) through a newly enabled control socket (sovrn_litestream_socket: /var/run/litestream.sock).-waitblocks until remote replication completes — exit 0 proves R2 reachable AND caught up. Daemon down, socket missing, R2 unreachable, or a missing DB all fail loudly and the marker goes stale. Marker max age 15m.backup-rclone): the existing 15-min oneshot unit now ends withtouch+chmod 0644ofrclone.ok— reached only when both syncs succeed. Marker max age 45m.{{ sovrn_datadir }}/health(root:sovrn 0750, created by the sovrnd role; markers mtime-only 0644, no secrets).recovery-bootstrapenables the freshness timer and runs both oneshots once after services start, so the health gate sees current markers instead of waiting up to 15 min.healthz changes
backupcheck split intobackup-litestream/backup-rclone(7 checks total; staged-failure matrix updated). Config keys are nowhealth.litestreammarkerfile/maxage+health.rclonemarkerfile/maxage(viper defaults 15m/45m; zero-config backstops innewHealthHandler).Verification (all green)
go vet ./...,SOVRN_INTEGRATION=0 go test ./...(full suite), gofmt clean on touched Go files.bash -n+shellcheck -S warningon all three sync scripts; edited Ansible YAML parses.status -jsonreflects only local state (reportsokeven with an unreachable replica — NOT a valid freshness probe), which is why the poller usessync -waitinstead.