Files
oikos/.agents/skills/service-health-check/SKILL.md
dtoro 5e3b946ded cleanup: fix all stale references across .agents/ docs
- Rewrite AGENTS.md: DB as source of truth, MCP knowledge tools, archive refs
- Fix OIKOS.md: seeds/ paths, remove Python-era notes, update deployment status
- Fix commands.md, agent-enrollment.md: archive/knowledge/ links
- Fix all SKILL.md files: remove hosts/*.yaml refs, point to inventory.yaml
- Fix HERMES.md, schema.md, page-templates.md, llm-wiki.md: update paths
- Fix bootstrap.sh: identity check reads inventory.yaml
- Fix README.md, cutover-checklist.md: stale wiki references
- Move convert-wiki.py to archive/ (one-shot done)
2026-07-07 21:00:39 +02:00

1.3 KiB

name, risk_class, inputs, verification, docs_update_checklist
name risk_class inputs verification docs_update_checklist
service-health-check read_only
service_name
homelab service <name> health

Service health check

Goal: determine whether a service is actually healthy, without ad-hoc SSH.

  1. homelab service <name> explain — read the context card: backend, blast radius, doc pointer, risk notes.
  2. homelab service <name> health — live health probe (HTTP code against the service's url/endpoint). Once the Week-3 scheduler ships, this reads a cached snapshot by default; pass --live to force a fresh probe.
  3. If unhealthy, homelab service <name> log (or MCP tail_log) for the last 200 lines.
  4. Cross-check blast radius: homelab node <name> relations — is this entity's own backend host healthy? A downstream failure (e.g. strong down) will show up here before the service's own logs explain anything.
  5. If the fix is a restart: classify first (seeds/policy.yamlservice-restart is reversible_low unless the service has a service_overrides entry, e.g. caddy/dns are config_mutation). Unattended agents may act on reversible_low without approval.

Docs-update checklist: none for a pure health check. If the investigation reveals stale risk_notes or a wrong doc_page, fix inventory.yaml in the same session.