- Rewrite AGENTS.md: DB as source of truth, MCP knowledge tools, archive refs - Fix OIKOS.md: seeds/ paths, remove Python-era notes, update deployment status - Fix commands.md, agent-enrollment.md: archive/knowledge/ links - Fix all SKILL.md files: remove hosts/*.yaml refs, point to inventory.yaml - Fix HERMES.md, schema.md, page-templates.md, llm-wiki.md: update paths - Fix bootstrap.sh: identity check reads inventory.yaml - Fix README.md, cutover-checklist.md: stale wiki references - Move convert-wiki.py to archive/ (one-shot done)
31 lines
1.3 KiB
Markdown
31 lines
1.3 KiB
Markdown
---
|
|
name: service-health-check
|
|
risk_class: read_only
|
|
inputs: [service_name]
|
|
verification: "homelab service <name> health"
|
|
docs_update_checklist: []
|
|
---
|
|
|
|
# Service health check
|
|
|
|
Goal: determine whether a service is actually healthy, without ad-hoc SSH.
|
|
|
|
1. `homelab service <name> explain` — read the context card: backend,
|
|
blast radius, doc pointer, risk notes.
|
|
2. `homelab service <name> health` — live health probe (HTTP code against
|
|
the service's `url`/`endpoint`). Once the Week-3 scheduler ships, this
|
|
reads a cached snapshot by default; pass `--live` to force a fresh probe.
|
|
3. If unhealthy, `homelab service <name> log` (or MCP `tail_log`) for the
|
|
last 200 lines.
|
|
4. Cross-check blast radius: `homelab node <name> relations` — is this
|
|
entity's own backend host healthy? A downstream failure (e.g. `strong`
|
|
down) will show up here before the service's own logs explain anything.
|
|
5. If the fix is a restart: classify first (`seeds/policy.yaml` —
|
|
`service-restart` is `reversible_low` unless the service has a
|
|
`service_overrides` entry, e.g. `caddy`/`dns` are `config_mutation`).
|
|
Unattended agents may act on `reversible_low` without approval.
|
|
|
|
Docs-update checklist: none for a pure health check. If the investigation
|
|
reveals stale `risk_notes` or a wrong `doc_page`, fix `inventory.yaml` in
|
|
the same session.
|