Files
oikos/README.md
dtoro d90de0759c docs: redesign README for agent clarity and usability
Problem: README.md was human-centric and lacked critical context for agents
(LLMs running on enrolled homelab clients). Agents needed:
- Explicit entry points (AGENTS.md → OIKOS.md → skills → MCP/files)
- Decision tree for tool selection (when to use MCP vs files vs grep)
- Explanation of operating model (OODA loop, risk classes, layer model)

Solution: Reorganized README with agent-first sections while preserving existing
human-useful content:

NEW SECTIONS:
- "For Agents" (entry points + MCP tool selection table with decision criteria)
- "Understanding the Operating Model" (Mermaid OODA loop diagram, risk classes,
  decision flow: classify → escalate if needed → execute → document)
- "Finding & Understanding Information" (layer model table: sources/wiki/index/log,
  what's immutable vs editable, when to update docs)

REVISED SECTIONS:
- "Map & Quick Navigation" (agent entry points first, then topology)
- "Conventions" (expanded with agent-specific guidance: caveman.md, page-templates.md)
- "Updating the Wiki" (clarified infrastructure changes vs restructuring;
  reinforced same-session update rule with explicit checklist)
- "More Information" (grouped agent-facing resources: HERMES, operations,
  skills, shared conventions)

All links verified. No new files needed — all referenced content already exists.

Verification:
- OODA loop diagram present (visual roadmap for decision flow)
- MCP vs Files vs Shell table shows decision criteria
- Layer model (sources/wiki/index/log) explained with immutability matrix
- All cross-references resolve
- Existing topology + infrastructure content preserved

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 18:32:08 +02:00

16 KiB
Raw Blame History

Homelab Wiki — hubris

Living documentation for the hubris Proxmox homelab + Oikos operating system.

For agents running on enrolled clients: start with AGENTS.md, then OIKOS.md.

Last refreshed against live state: 2026-07-06.


For Agents — Navigation & Entry Points

You are running on a client enrolled in the hubris homelab

  1. First: Read AGENTS.md once. It explains who you are, the topology, available tools, conventions, and how to act.
  2. Before any mutation: Read OIKOS.md. It defines the operating model, risk classes, approval flow, and the ontology you'll consult.
  3. For specific workflows: Load the matching skill from .agents/skills/<name>/SKILL.md (e.g., service-health-check).
  4. When in doubt: Use MCP tools (search_docs, get_page, explain, get_changelog) — they're cheaper and more reliable than grepping.

Key References for Agents

  • What am I?/opt/homelab-context/hosts/<hostname>.yaml (read on first run)
  • Live topologyinventory.yaml + hosts/*.yaml (canonical, always wins)
  • Risk & approvaloikos/policy.yaml (enforced, not advisory)
  • Runbooks & workflows.agents/skills/ (risk class + verification checklist included)
  • State of OikosOIKOS.md build status (scheduled probes, drift detectors, signals, approval engine)

When to Use MCP vs Files vs Shell

Task Use Tool
Resolve hostname → address MCP get_host(name) or list_services()
Search wiki by content MCP search_docs(query)
Read a wiki page MCP or file get_page(path) or cat knowledge/wiki/.../...md
Get changelog entries MCP get_changelog(page, since?)
Understand a service MCP explain(service) — compact context card, cheaper than search+read
Blast-radius query MCP get_relations(entity) (ontology walk)
List available secrets MCP list_my_secrets() (scoped to your age key)
Browse or grep File Raw grep when MCP unreachable, or exploratory browsing

When MCP is unreachable: fall back to grepping the clone at /opt/homelab-context/. The local files are the same; MCP is just an index.


Understanding the Operating Model

Before you act, classify your action against oikos/policy.yaml.

The Oikos OODA Loop + Decision Tree

┌─────────────────────────────────────────────────────────────┐
│                    Observe                                  │
│        (probes, drift detectors, agent signals)             │
└────────────────────┬────────────────────────────────────────┘
                     │
┌────────────────────▼────────────────────────────────────────┐
│                    Orient                                   │
│      (ontology, context, state, entity relations)           │
└────────────────────┬────────────────────────────────────────┘
                     │
┌────────────────────▼────────────────────────────────────────┐
│  Decide: Classify against oikos/policy.yaml                │
│                                                              │
│  ├─ read_only                → [unattended, no ledger]      │
│  ├─ reversible_low           → [unattended + ledger entry]  │
│  ├─ config_mutation          → [ESCALATE: operator approval]│
│  └─ destructive              → [ESCALATE + confirmation]    │
└────────────────────┬────────────────────────────────────────┘
                     │
        ┌────────────┴────────────┐
        │                         │
   ┌────▼─────┐            ┌──────▼──────┐
   │ Auto-act │            │  Escalate   │
   └────┬─────┘            │ (get approval
        │                  │  via homelab │
┌───────▼──────────────────┴──────────┐  │
│            Act                       │  │
│  (homelab CLI, runbooks, skills)    │  │
└────────────────┬─────────────────────┘  │
                 │                        │
            ┌────▼────────────────────────┘
            │
┌───────────▼──────────────────────────────┐
│          Verify                           │
│    (checklist from SKILL.md)              │
└───────────┬──────────────────────────────┘
            │
┌───────────▼──────────────────────────────┐
│          Ledger                           │
│     (mutation record: who/what/risk)      │
└───────────┬──────────────────────────────┘
            │
┌───────────▼──────────────────────────────┐
│        Document                           │
│   (wiki update, same-session rule)        │
└───────────┬──────────────────────────────┘
            │
            └─────────────────┐
                              │
                    ┌─────────▼─────────┐
                    │ Loop back to      │
                    │    Observe        │
                    └───────────────────┘

Risk Classes (enforced, not advisory)

From oikos/policy.yaml:

  • read_only — status, logs, docs, inventory queries. Unattended. MCP tools are all read_only.
  • reversible_low — restart, cache clear, sync pull. Unattended + ledger entry.
  • config_mutation — tracked-config edits (commit+push, never local), deploys, upgrades, DNS/ingress changes. Operator approval required.
  • destructive — destroy, format, wipe, rotate, revoke. Approval + typed confirmation phrase.

Decision Flow

  1. Decide: Use homelab decide <action> <entity> to classify (risk class × blast radius × confidence).
  2. Escalate if needed: homelab approval request (Matrix-delivered to operator; see operations/commands.md).
  3. Execute: Use homelab CLI (not ad-hoc SSH) — it enforces policy, logs mutations, and verifies outcomes.
  4. Document: Update wiki in the same session (per AGENTS.md §5 and the same-session rule).

The Ontology Graph

Everything that can break, be changed, or hold data has an entity in inventory.yaml + oikos/ontology.yaml. Blast-radius questions ("what breaks if strong goes down?") are graph walks via homelab node <name> relations, not doc archaeology.

See: OIKOS.md (full operating model, OODA loop, primitives, lifecycle gates, build status).


Finding & Understanding Information

The narrative documentation is organized in layers:

Layer What it is Where Immutable? How agents use it
Sources Raw evidence: incidents, external refs, live state knowledge/sources/investigations/ Yes Read to understand root causes; do not rewrite
Wiki Synthesized current-state: one page per node & per system knowledge/wiki/{containers,hosts,vms,infrastructure}/ No This is the reference layer — if wiki disagrees with live state, update it in the same session
Index Pure listings — every page in scope with one-line summary index.md / folder README.md No Navigation aid; keep it current when wiki restructures
Log Append-only doc-maintenance record (restructures, ingests, lints) knowledge/log.md Yes (append-only) Read to understand past doc changes; never edit directly

Changelog ≠ Log: Each wiki page ends with a ## Changelog (infrastructure changes to that node, machine-parsed). That's not the Log; the Log records doc operations only.

See: llm-wiki.md (full rules, page structure, immutability contract).


Map & Quick Navigation

Agent Entry Points (Start Here)

Topology & Infrastructure

Proxmox Hosts

VMs

LXC Containers

See the full table with IPs, hosts, mounts, and status in containers/index.md. Quick summary:

  • hubris (15 active): 102 nfs-export, 103 paperless, 104 gitea, 105 apps, 114 nextcloud, 119 sophia, 120 mule-images, 121 caddy, 124 authentik, 128 trmnl, 132 rclone
  • strong (7 active): 101 jellyfin, 118 elementsynapse, 122 arriman, 129 house, 130 grimmory, 133 seanime, 134 romm
  • Destroyed (archaeology): see containers/index.md

Cross-Cutting Infrastructure

Knowledge & References


Conventions

All pages follow:

  • File naming. Foundational docs (entry-points, agent instruction, references) are ALL-CAPS (AGENTS.md, OIKOS.md, GLOSSARY.md); containers use <id>-<name>.md; infrastructure pages use lowercase-with-dashes; plans and incidents use YYYY-MM-DD-slug.md; skills are <name>/SKILL.md. See page-templates.md for the full rules.
  • Voice & vocabulary. Concise, technical, sysadmin-to-sysadmin. No marketing prose, no puffers (seamless, robust, leverage, etc.). Full rules in writing-style.md.
  • Cross-linking is mandatory. If a page references a node or system, link to it. Treat orphans as a bug.
  • Live state wins. When something here disagrees with pct config / docker inspect / running state, fix the wiki and add a changelog entry in the same session.
  • Tracked configs. Pages for configs living in git repos (Caddy, Gitea, Artifacto, mule-image) must note the repo. Edits go through commit+push, never local changes. See auto-deploy.
  • No secrets. This is a private repo, but still: reference secret paths, never secret values.

For agents: Read caveman.md (terse communication standard). Use templates at page-templates.md when creating pages.


Updating the Wiki

When You Change Infrastructure

  1. Update the relevant page (config snapshot, ports, mounts, IP address).
  2. Add a ### YYYY-MM-DD — title entry to the page's ## Changelog section (reverse chronological order).
  3. If the change touches a cross-cutting system (DNS, Caddy, Authentik, mesh), update that page too and link from the changelog.
  4. If it's an incident, add a record to knowledge/sources/investigations/.

When You Restructure the Wiki

  1. Update the relevant index.md / README.md in that section.
  2. Add a single-line entry to knowledge/log.md: ## [YYYY-MM-DD] <operation> | <summary> (e.g., ## [2026-07-06] restructure | split infrastructure/dns into dns.md + dns-advanced.md).

The Same-Session Update Rule

Any meaningful state change made in this session requires a wiki update before the session closes. A change that touches a container page must also update:

  • The containers/index.md table (IPs, host, mounts, status)
  • The root README.md table (if affected)
  • The Caddy page site list (if affects *.hubris.network routing)
  • The DNS / ingress infrastructure pages (if affects routing)
  • The hosts/hubris.md or hosts/strong.md page (if container count changes)
  • The inventory.yaml host entry (source of truth for hosts/*.yaml generation)
  • The knowledge/wiki/infrastructure/topology.md (regenerate if needed)

Not updating all linked places is a bug. See page-templates.md — same-session update rule.


More Information

  • For Hermes agentsHERMES.md (persona, source-of-truth hierarchy, token efficiency)
  • For manual workflows.agents/operations/ (commands cheatsheet, agent enrollment, Hermes guide)
  • For skills/runbooks.agents/skills/ (load the matching SKILL.md before acting; includes risk class + verification)
  • MCP toolsAGENTS.md §3 (available tools, when to use MCP vs files)
  • Page templates & voice.agents/shared/ (page-templates.md, writing-style.md, caveman.md, llm-wiki.md)
  • Machine-readable substrateinventory.yaml, oikos/policy.yaml, oikos/ontology.yaml (not part of the wiki; see llm-wiki.md)