Files
oikos/.agents/skills/incident-investigation/SKILL.md
dtoro 5c5016b3c7 docs: reshape runbooks into .agents/skills/<name>/SKILL.md (phase 4)
Problem: runbooks are agent-executable procedures but lived at the repo root,
separate from the other agent instruction now under .agents/.

Change:
- Move runbooks/<name>.md -> .agents/skills/<name>/SKILL.md (folder per skill,
  matching the wiki-hq skills layout). Frontmatter (name, risk_class, inputs,
  verification, docs_update_checklist, transition) preserved.
- Rewrite links (inbound from plans; between-skill siblings) via the move map.
- Update prose references in AGENTS.md, HERMES.md, .agents/OIKOS.md, and the
  operations schema; fix a pre-existing stale link to operations/commands.md.

No code consumed runbooks/ by path, so nothing else changes.

Verification: all SKILL.md frontmatter parses with valid risk_class; every
lifecycle transition resolves to an oikos/ontology.yaml state; broken-link
count 127 -> 126 (fixed one, introduced none).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:39:31 +02:00

1.6 KiB

name, risk_class, inputs, verification, docs_update_checklist
name risk_class inputs verification docs_update_checklist
incident-investigation read_only
symptom
affected_entity
n/a — investigation produces a written record, not a state change
investigations_entry

Incident investigation

Goal: understand what broke and why, before touching anything.

  1. homelab service <name> explain (or homelab node <name> relations if the affected entity is a host) — get the blast radius and doc pointer first. Don't start pulling logs blind.
  2. homelab service <name> health + homelab service <name> log (or MCP get_service_status / tail_log) for the affected service.
  3. Walk the blast radius: is a shared dependency down (caddy, dns, authentik, or the backend host itself)? homelab node <name> relations shows "affected by" — check those first.
  4. homelab apt-audit if the symptom looks like a dpkg/upgrade interaction.
  5. Check the change ledger for recent mutations to the affected entity or anything upstream of it: homelab service <name> history (once populated) or grep ledger/*.jsonl.
  6. Write findings to a new investigations/<date>-<slug>.md — symptom, timeline, root cause, fix applied, prevention. This is the durable record; don't rely on chat history.

Docs-update checklist: always create the investigation entry. If the root cause was stale/wrong inventory data (a doc_page, config_repo, or backend that didn't match reality — this happened during Week 1 kernel work, see the authentik backend fix), correct inventory.yaml in the same session.