Files
oikos/runbooks/incident-investigation.md
dtoro f6b57cbe3a Oikos Week 2: Service Console v0, change ledger, node relations, runbooks
Adds the shared kernel modules (oikos/policy.py, oikos/relations.py,
oikos/ledger.py) that let every surface — CLI, MCP, context-card
generator — agree on risk classification and ontology graph walks
from one implementation.

homelab CLI: `service <name> explain|health|docs|log|actions|history`
(Service Console v0), `change preflight <service>`, `node <name>
relations`. Restart and client add/remove now append change-ledger
entries (ledger/*.jsonl, committed alongside the change they record).

mcp/server.py mirrors explain/preflight/get_relations/get_change_history
as MCP tools, card-first so agent orientation is one call instead of
several search_docs/get_page round-trips.

oikos/gen-topology.py now also emits a compact context card per host
and service (oikos/cards/*.md) — identity, blast radius, safe actions +
risk class, doc pointer, recent ledger history.

runbooks/*.md: service health check, config change + deploy, client
enrollment, incident investigation, and the five node lifecycle
transitions (provision/activate/migrate/deprecate/destroy), each with
machine-readable frontmatter (risk class, inputs, verification,
docs-update checklist). Wired into HERMES.md so agents load these
instead of rediscovering topology per-task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 23:02:32 +02:00

1.6 KiB

name, risk_class, inputs, verification, docs_update_checklist
name risk_class inputs verification docs_update_checklist
incident-investigation read_only
symptom
affected_entity
n/a — investigation produces a written record, not a state change
investigations_entry

Incident investigation

Goal: understand what broke and why, before touching anything.

  1. homelab service <name> explain (or homelab node <name> relations if the affected entity is a host) — get the blast radius and doc pointer first. Don't start pulling logs blind.
  2. homelab service <name> health + homelab service <name> log (or MCP get_service_status / tail_log) for the affected service.
  3. Walk the blast radius: is a shared dependency down (caddy, dns, authentik, or the backend host itself)? homelab node <name> relations shows "affected by" — check those first.
  4. homelab apt-audit if the symptom looks like a dpkg/upgrade interaction.
  5. Check the change ledger for recent mutations to the affected entity or anything upstream of it: homelab service <name> history (once populated) or grep ledger/*.jsonl.
  6. Write findings to a new investigations/<date>-<slug>.md — symptom, timeline, root cause, fix applied, prevention. This is the durable record; don't rely on chat history.

Docs-update checklist: always create the investigation entry. If the root cause was stale/wrong inventory data (a doc_page, config_repo, or backend that didn't match reality — this happened during Week 1 kernel work, see the authentik backend fix), correct inventory.yaml in the same session.