T6: consolidated cell healthz in sovrnd

closed
#b957ddb opened by agent Sep 13

Parent: bug 75966cc (cell architecture tracking). Gates floating-IP moves; keep metrics out (see metrics issue).

Goal

One endpoint expressing whole-cell health, suitable for paging and as the pre-move gate in the floating-IP recovery runbook.

Scope

  • GET /healthz/ready (+ /live): aggregate with per-check timeouts and cached results — Stalwart ready + authed JMAP probe, ZDS describeServer, Unbound test resolutions (cell MX + external name), SQLite writability, disk + Litestream/rclone backup freshness.
  • JSON per-component body for humans, binary status code for automation.
  • Unit + integration tests; thresholds/check list documented.

Acceptance

  • Single endpoint reflects cell state truthfully under staged failures (each dependency down => non-OK with identifying body); docs merged.

3 Comments

agent bf93567 Sep 16

Consolidated GET /healthz Implementation Plan (bug b957ddb)

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Add unauthenticated GET /healthz to sovrnd returning 200 + per-check JSON when healthy, 503 + same JSON shape when any dependency fails.

Architecture: New healthz.go (package sovrn, next to caddy_ask.go) with an injectable checker set run concurrently under per-check timeouts; mounted on the unauthenticated top mux in router.go (not the authed XRPC mux). No caching in v1 — concurrency + tight timeouts keep it under Larm’s 10s probe timeout; Larm’s multi-location majority voting means the handler must be concurrency-safe.

Tech Stack: Go stdlib net/http + http.ServeMux, existing stalwart.Client (JMAP), store.Store (SQLite), domain.Resolver DNS, viper config, httptest tests.

Decisions locked: path GET /healthz only (not /healthz/ready); 503 on any failure; unauthenticated public; endpoint-only scope (Larm monitor configured manually — recommended settings documented in Task 5).

Larm finding (https://docs.larm.dev/monitors): HTTP monitor, Method: GET, URL: https://<cell>/healthz, Expected status codes: [200] — status code is the sole error signal (no keyword needed; body shape identical on success/failure). Timeout: 10s, Confirm down: 1 min, Confirm up: 3 min defaults; no automated recovery — alerts only, operator runs docs/runbooks/recovery-primary-ip.md.


Task 1: Health response types + handler skeleton

Files: - Create: healthz.go - Test: healthz_test.go

  • [ ] Step 1: Write the failing test (TestHealthzAllGreen — single ok checker → 200 + status: ok JSON).
  • [ ] Step 2: Run test to verify it fails (go test -run TestHealthzAllGreen ., expect FAIL).
  • [ ] Step 3: Write minimal implementation (checkResult, healthResponse, checkFunc, healthHandler.ServeHTTP with concurrent per-check timeouts, 200⁄503 aggregation, application/json).
  • [ ] Step 4: Run test to verify it passes (go test -run TestHealthzAllGreen ., expect PASS).

Task 2: Failure semantics — any fail → 503, same body

Files: - Modify: healthz_test.go (staged-failure matrix over stalwart/zds/dns/sqlite/disk/backup)

  • [ ] Step 1: Write TestHealthzAnyFailureIs503 (each single failing check → 503 + status: degraded + identifying error).
  • [ ] Step 2: Run, expect PASS against Task 1 skeleton.

Task 3: Real checkers — Stalwart, ZDS, DNS, SQLite, disk, backup

Files: - Modify: healthz.go (constructors), healthz_test.go (fake-server tests), config.go (+ HealthConfig, defaults, bindEnv), config_test.go, deployment/roles/sovrnd/templates/sovrn.toml.j2 ([health] block)

  • [ ] Checker contracts (each honors ctx timeout): stalwart = GET {url}/healthz/ready + authed JMAP Query(Tenant); zds = GET {zdsurl}/xrpc/com.atproto.server.describeServer; dns = resolve cell MX + external name via cell Unbound; sqlite = writability probe (SELECT + transactional write/rollback, no residue); disk = statfs used% vs threshold (default 85); backup = Litestream/rclone marker mtime vs max age (default 10m).
  • [ ] Config: HealthConfig{PerCheckTimeout 2s, ZDSURL, CellMX, ExternalName example.com, DiskPath=DataDir, DiskMaxUsedPct 85, BackupMaxAge 10m, LitestreamState}.

Task 4: Wire GET /healthz into the router (unauthenticated, public)

Files: - Modify: router.go (top.Handle("GET /healthz", logging.Middleware(newHealthHandler(cfg, st, sw), nil)) before / catch-all), router test (open endpoint, JSON content-type, 200-or-503), sovrn.toml.j2 + inventory vars.

Task 5: Docs — check list, thresholds, Larm monitor recipe

Files: - Modify: docs/deployment.md, docs/runbooks/provision-cell.md, docs/runbooks/recovery-primary-ip.md, sovrn.toml.j2 comments.

Document each check, failure meaning, defaults, and the Larm recipe (URL, GET, expected [200], confirm down 1m/up 3m, alert channels; no recovery webhook). Final: go test ./... + go vet ./... green.

agent b49e5f7 Sep 16

Implementation complete: consolidated GET /healthz

Implemented per the plan above (comment bf93567). All in the working copy, uncommitted (the tree already held unrelated backup/Caddy work — left untouched).

What was built

  • healthz.go (new): unauthenticated GET /healthz, 200 + {"status":"ok",...} when healthy, 503 + {"status":"degraded",...} with the same per-check shape otherwise. Six checks, run concurrently under health.perchecktimeout (2s):
    • stalwart: GET {stalwart.url}/healthz/ready + authed JMAP x:Tenant/query probe (either leg names itself on failure)
    • zds: every active PDS instance from the registry, describeServer on its loopback port with origin verification (supervisor’s healthGate semantics). Supervision OFF skips ok (caddy-ask rationale); no actives skips ok. No static URL — ZDS is per-domain instances, so the planned health.zdsurl was dropped.
    • dns: resolves configured cell MX + external names via system resolver (Unbound on cells); names failures
    • sqlite: Store.Ping — SELECT plus rolled-back INSERT (no residue)
    • disk: statfs used% vs threshold (default 85)
    • backup: marker-file mtime vs max age (default 10m); empty marker skips “unconfigured” ok until the sync scripts touch it (follow-up, noted in provision-cell.md)
  • router.go: mounted on the unauthenticated top mux before the / catch-all; Caddy needs no change (only /internal/* is blocked at the edge).
  • config.go + sovrn.toml.j2 + sovrn.yml: new [health] section (perchecktimeout, cellmx defaulting to mail hostname, externalname default example.com, diskmaxusedpct, backupmaxage, backupmarkerfile).
  • internal/store: Ping added to the Store interface (sole impl is sqlite).
  • Docs: deployment.md (check list, thresholds, Larm recipe), both runbooks now gate on GET /healthz instead of the interim manual checks.

Larm setup (manual, endpoint-only scope)

HTTP monitor: GET https://<cell>/healthz, expected status [200], timeout 10s, confirm down 1m / up 3m. Status code is the sole signal — no keyword matching (passing checks also contain “ok”). Alerts only; recovery stays manual per the runbook.

Verification

  • 13 new tests: all-green, 405, six-way staged-failure matrix (each single failure → 503 + identifying body), per-checker fakes, supervised-registry path, zero-config defaults, end-to-end via NewApp (503 naming dead Stalwart).
  • go vet ./... clean; SOVRN_INTEGRATION=0 go test ./... all packages pass; root package also passes under -race.

Leaving open for your review — close when satisfied.

agent b891547 Sep 16

Backup markers wired: litestream poller + rclone touch

The backup leg is now two live checks instead of “unconfigured”.

How freshness is proven

  • Litestream (backup-litestream): continuous replicate has no periodic hook, so a new 5-min timer (litestream-freshness.timer, sovrn_litestream_freshness_schedule: "*:0/5") runs litestream-freshness.sh, which executes litestream sync -wait -timeout 30 <db> per database (stalwart.db, sovrn.db, oauth.db, ZDS *.db) through a newly enabled control socket (sovrn_litestream_socket: /var/run/litestream.sock). -wait blocks until remote replication completes — exit 0 proves R2 reachable AND caught up. Daemon down, socket missing, R2 unreachable, or a missing DB all fail loudly and the marker goes stale. Marker max age 15m.
  • rclone (backup-rclone): the existing 15-min oneshot unit now ends with touch + chmod 0644 of rclone.ok — reached only when both syncs succeed. Marker max age 45m.
  • Markers live in {{ sovrn_datadir }}/health (root:sovrn 0750, created by the sovrnd role; markers mtime-only 0644, no secrets).
  • Recovery: recovery-bootstrap enables the freshness timer and runs both oneshots once after services start, so the health gate sees current markers instead of waiting up to 15 min.

healthz changes

  • Single backup check split into backup-litestream / backup-rclone (7 checks total; staged-failure matrix updated). Config keys are now health.litestreammarkerfile/maxage + health.rclonemarkerfile/maxage (viper defaults 15m/45m; zero-config backstops in newHealthHandler).

Verification (all green)

  • go vet ./..., SOVRN_INTEGRATION=0 go test ./... (full suite), gofmt clean on touched Go files.
  • bash -n + shellcheck -S warning on all three sync scripts; edited Ansible YAML parses.
  • Investigated the real v0.5.13 binary first: status -json reflects only local state (reports ok even with an unreachable replica — NOT a valid freshness probe), which is why the poller uses sync -wait instead.
  • Freshness script tested end-to-end against a live daemon + file replica: all-sync touches the marker; daemon down fails exit 1 with no marker; missing DB fails naming the file.