Files
oikos/runbooks/service-health-check.md
dtoro f6b57cbe3a Oikos Week 2: Service Console v0, change ledger, node relations, runbooks
Adds the shared kernel modules (oikos/policy.py, oikos/relations.py,
oikos/ledger.py) that let every surface — CLI, MCP, context-card
generator — agree on risk classification and ontology graph walks
from one implementation.

homelab CLI: `service <name> explain|health|docs|log|actions|history`
(Service Console v0), `change preflight <service>`, `node <name>
relations`. Restart and client add/remove now append change-ledger
entries (ledger/*.jsonl, committed alongside the change they record).

mcp/server.py mirrors explain/preflight/get_relations/get_change_history
as MCP tools, card-first so agent orientation is one call instead of
several search_docs/get_page round-trips.

oikos/gen-topology.py now also emits a compact context card per host
and service (oikos/cards/*.md) — identity, blast radius, safe actions +
risk class, doc pointer, recent ledger history.

runbooks/*.md: service health check, config change + deploy, client
enrollment, incident investigation, and the five node lifecycle
transitions (provision/activate/migrate/deprecate/destroy), each with
machine-readable frontmatter (risk class, inputs, verification,
docs-update checklist). Wired into HERMES.md so agents load these
instead of rediscovering topology per-task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 23:02:32 +02:00

1.3 KiB

name, risk_class, inputs, verification, docs_update_checklist
name risk_class inputs verification docs_update_checklist
service-health-check read_only
service_name
homelab service <name> health

Service health check

Goal: determine whether a service is actually healthy, without ad-hoc SSH.

  1. homelab service <name> explain — read the context card: backend, blast radius, doc pointer, risk notes.
  2. homelab service <name> health — live health probe (HTTP code against the service's url/endpoint). Once the Week-3 scheduler ships, this reads a cached snapshot by default; pass --live to force a fresh probe.
  3. If unhealthy, homelab service <name> log (or MCP tail_log) for the last 200 lines.
  4. Cross-check blast radius: homelab node <name> relations — is this entity's own backend host healthy? A downstream failure (e.g. strong down) will show up here before the service's own logs explain anything.
  5. If the fix is a restart: classify first (oikos/policy.yamlservice-restart is reversible_low unless the service has a service_overrides entry, e.g. caddy/dns are config_mutation). Unattended agents may act on reversible_low without approval.

Docs-update checklist: none for a pure health check. If the investigation reveals stale risk_notes or a wrong doc_page, fix inventory.yaml in the same session.