monitor

Health-endpoint monitor on Cloudflare Workers + Durable Objects. See DESIGN.md for the brief.

  • One MonitorDO per region (wnam, enam, weur, eeur, apac), placed with location hints. Each runs an alarm every CHECK_INTERVAL_SECONDS, GETs every endpoint, and stores results in SQLite (a trigger drops rows older than 14 days).
  • Alerts go to ALERT_TO when an endpoint starts failing in a region, every REMINDER_HOURS while it keeps failing, and when it recovers. Anything other than HTTP 200 (including timeouts) is a failure.
  • Each alert also goes to Pushover once its secrets are set: a second, independent channel. Each channel tracks its own last alert, so a failed send is retried on that channel only. Pushover messages list the failing entries of a /healthz checks object; recoveries arrive at priority -1 (quiet).
  • An hourly cron (and any homepage visit) makes sure every region’s alarm loop is running.
  • https://monitor.rtw.run/ shows the latest result per endpoint; /endpoint?url=… shows history.

Configuration

Endpoints, interval, timeout and alert settings are vars in wrangler.jsonc; change them and redeploy. Basic-auth credentials are secrets (all requests are rejected until both are set):

npx wrangler secret put AUTH_USER
npx wrangler secret put AUTH_PASS

Pushover is off until both of its secrets are set (an application token and your user key):

npx wrangler secret put PUSHOVER_TOKEN
npx wrangler secret put PUSHOVER_USER

Development

pnpm test        # vitest in the Workers runtime
pnpm typecheck
pnpm dev         # local server; credentials come from .dev.vars (AUTH_USER=…, AUTH_PASS=…)
pnpm deploy
pnpm types       # regenerate worker-configuration.d.ts after editing wrangler.jsonc