--- name: service-health-check risk_class: read_only inputs: [service_name] verification: "MCP get_service_status" docs_update_checklist: [] --- # Service health check Goal: determine whether a service is actually healthy, without ad-hoc SSH. 1. MCP `explain` — read the context card: backend, blast radius, doc pointer, risk notes. 2. MCP `get_service_status` — live health probe (HTTP code against the service's `url`/`endpoint`); the scheduler also probes on its own interval, so this may reflect a recent cached result, not necessarily a fresh one. 3. If unhealthy, `tail_log` for the last 200 lines. 4. Cross-check blast radius: MCP `get_blast_radius` — is this entity's own backend host healthy? A downstream failure (e.g. a Proxmox host down) will show up here before the service's own logs explain anything. 5. If the fix is a restart: classify first (`seeds/policy.yaml` — `service-restart` is `reversible_low` unless the service has a `service_overrides` entry, e.g. `caddy`/`dns` are `config_mutation`). Unattended agents may act on `reversible_low` without approval. Docs-update checklist: none for a pure health check. If the investigation reveals stale `risk_notes` or a wrong `doc_page`, fix `inventory.yaml` in the same session.