Problem: runbooks are agent-executable procedures but lived at the repo root, separate from the other agent instruction now under .agents/. Change: - Move runbooks/<name>.md -> .agents/skills/<name>/SKILL.md (folder per skill, matching the wiki-hq skills layout). Frontmatter (name, risk_class, inputs, verification, docs_update_checklist, transition) preserved. - Rewrite links (inbound from plans; between-skill siblings) via the move map. - Update prose references in AGENTS.md, HERMES.md, .agents/OIKOS.md, and the operations schema; fix a pre-existing stale link to operations/commands.md. No code consumed runbooks/ by path, so nothing else changes. Verification: all SKILL.md frontmatter parses with valid risk_class; every lifecycle transition resolves to an oikos/ontology.yaml state; broken-link count 127 -> 126 (fixed one, introduced none). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1.6 KiB
1.6 KiB
name, risk_class, inputs, verification, docs_update_checklist
| name | risk_class | inputs | verification | docs_update_checklist | |||
|---|---|---|---|---|---|---|---|
| incident-investigation | read_only |
|
n/a — investigation produces a written record, not a state change |
|
Incident investigation
Goal: understand what broke and why, before touching anything.
homelab service <name> explain(orhomelab node <name> relationsif the affected entity is a host) — get the blast radius and doc pointer first. Don't start pulling logs blind.homelab service <name> health+homelab service <name> log(or MCPget_service_status/tail_log) for the affected service.- Walk the blast radius: is a shared dependency down (
caddy,dns,authentik, or the backend host itself)?homelab node <name> relationsshows "affected by" — check those first. homelab apt-auditif the symptom looks like a dpkg/upgrade interaction.- Check the change ledger for recent mutations to the affected entity
or anything upstream of it:
homelab service <name> history(once populated) or grepledger/*.jsonl. - Write findings to a new
investigations/<date>-<slug>.md— symptom, timeline, root cause, fix applied, prevention. This is the durable record; don't rely on chat history.
Docs-update checklist: always create the investigation entry. If the
root cause was stale/wrong inventory data (a doc_page, config_repo,
or backend that didn't match reality — this happened during Week 1
kernel work, see the authentik backend fix), correct inventory.yaml
in the same session.