-- 027_check_last_health.up.sql -- Aggregate an entity's health across its checks instead of last-writer-wins. -- -- runCheck wrote entity_status.health on every check completion, so an -- entity's health was simply whichever of its checks finished most recently. -- host:hubris has 6 checks, host:strong 6 — one failing probe alternating with -- five passing ones produced a permanent flap: 226 health.changed events for -- host:strong in a single hour, oscillating down/healthy, while the host was -- fine the whole time. -- -- On this fleet the trigger is a known false positive: the scheduler's network -- vantage point cannot ICMP host:strong, so its ping check fails while every -- ssh-script check succeeds. Under last-writer-wins that one probe was enough -- to declare the whole host down, twice a minute. -- -- Storing each check's own verdict lets entity health be derived as the worst -- current result across that entity's enabled checks — so a single failing -- probe degrades the entity honestly without erasing what the other five say, -- and a passing probe cannot mask a genuine failure elsewhere. ALTER TABLE check_defs ADD COLUMN IF NOT EXISTS last_health TEXT; COMMENT ON COLUMN check_defs.last_health IS 'This check''s own most recent verdict (healthy/degraded/down/unknown). entity_status.health is the worst of these across the target''s enabled checks.'; -- The aggregation reads every enabled check for one target on each completion. CREATE INDEX IF NOT EXISTS idx_check_defs_target_health ON check_defs (target_id) WHERE enabled AND target_id IS NOT NULL;