Files
oikos/.agents/OIKOS.md
dtoro 8a6422bd7d docs: move narrative wiki under knowledge/wiki/ (phase 3)
Problem: node and cross-cutting narratives lived at the repo root
(containers/, vms/, infrastructure/, host .md files), interleaved with the
machine-readable substrate.

Change:
- Move containers/ -> knowledge/wiki/containers/, vms/ -> knowledge/wiki/vms/,
  infrastructure/ -> knowledge/wiki/infrastructure/, hosts/{hubris,strong}.md ->
  knowledge/wiki/hosts/, infrastructure/references/ -> knowledge/sources/references/,
  GLOSSARY.md -> knowledge/GLOSSARY.md.
- Add knowledge/{index.md,log.md,sources/index.md} scaffolding.
- Rewrite all relative links repo-wide via a path-resolving mapper (inbound +
  outbound + between-moved-files), including .hermes/, runbooks, operations,
  investigations, plans, README, AGENTS.
- Repoint inventory.yaml doc_page fields and regenerate hosts/*.yaml (which
  embed doc_page); update oikos/gen-topology.py output path, candidate doc
  paths, and footer links; update code-comment doc paths.

Substrate untouched in place: inventory.yaml, hosts/*.yaml (regenerated,
idempotent), oikos/ code, mcp/, secrets/, bin/.

Verification:
- Logical broken-link set identical to pre-move baseline (net 128 -> 127; the
  topology regen fixed one, introduced none). Remaining are pre-existing refs
  to destroyed/archived nodes, out of scope for this move.
- gen-topology.py --check exit 0 (in sync); cards carry knowledge/wiki/ doc paths.
- build_host_files.py idempotent; all inventory doc_page targets resolve.
- MCP contract verified: get_page/search_docs/get_changelog resolve moved pages.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:35:23 +02:00

13 KiB
Raw Blame History

Oikos — the operating model

Oikos (Greek: household) is the agent operating system layered on this repo. It is not new infrastructure: inventory.yaml is the kernel data structure, the homelab CLI and MCP server are the syscall surface, and this page defines the rules everything above them follows.

Read this after AGENTS.md. Machine-readable companions: oikos/ontology.yaml (systems model), oikos/policy.yaml (risk & approval).

The kernel loop: OODA

Every Oikos activity — scheduled probe, agent task, operator request — is one pass through Observe → Orient → Decide → Act:

  1. Observe — probes, drift detectors, and agent findings produce Signals (structured records, not loose messages): pending updates, high temperature, low disk, service down, cert expiry, stale backup, inventory drift.
  2. Orient — walk the ontology graph: what entity is affected, what depends on it (blast radius), its lifecycle state, whether a runbook matches, what the ledger says about past attempts.
  3. Decide — the classifier scores risk class × blast radius × confidence and routes:
    • auto-act: within autonomy policy, high confidence, contained radius
    • escalate: operator approval via Matrix (/ reaction) or the Oikos Console's /approvals page (destructive actions additionally need a typed confirmation phrase either way)
    • queue: informational — console + reports The classifier can only lower autonomy relative to policy, never raise it. When in doubt, escalate.
  4. Act — execute through homelab commands or runbooks (never ad-hoc SSH), then verify with the action's verification command, write a ledger entry, resolve the Signal, and update docs in the same session.

Primitives

Primitive What it is Lives in
Host / Service topology entities inventory.yaml (+ generated hosts/*.yaml)
Secret SOPS+age encrypted value, per-client recipients secrets/ + .sops.yaml
Runbook executable workflow with risk class + verification runbooks/ (Week 2)
Signal something needing attention, with lifecycle signals/ ledger (Week 3)
Change one mutation: who, what, risk, approval, verification ledger/ (Week 2)
Approval short-TTL signed grant for a gated action approval engine (Week 3)
Incident investigation narrative investigations/
Plan design doc for non-trivial work plans/
Agent enrolled client identity = its age pubkey inventory.yaml + .sops.yaml

Risk classes (enforced, not advisory)

From oikos/policy.yaml:

  • read_only — status, logs, docs, inventory. Unattended.
  • reversible_low — restart, cache clear, sync pull. Unattended + ledger.
  • config_mutation — tracked-config edits (commit+push, never local), deploys, upgrades, DNS/ingress changes. Operator approval.
  • destructive — destroy, format, wipe, rotate, revoke. Approval + typed confirmation phrase.

Lifecycle gates modify these: provisioning nodes are freely mutable (nothing depends on them); deprecated nodes accept no new dependents; anything touching a destroyed node is drift.

The systems model

Eight domains — physical, compute, network, storage, software, identity & access, operations, external — cover everything in the lab; entities are connected by typed edges (hosts, provides, mounts, stores-on, routes-to, can-decrypt, depends-on, backs-up-to, …) defined in oikos/ontology.yaml. Rule of completeness: if it can break, be changed, or hold data, it has an entity and edges. Blast-radius questions ("what breaks if strong goes down?") are graph walks, not doc archaeology.

Nodes move through an explicit lifecycle — planned → provisioning → active → migrating → deprecated → destroyed — stored as state: in inventory (absent = active). Destroyed nodes live in the archaeology: section. Each transition is a runbook checklist; deprecation completes only when inbound edges reach zero.

Generated views: infrastructure/topology.md (Mermaid, regenerated from inventory) and the live, clickable version at oikos.hubris.network/graph once the Console is deployed.

Conventions carried forward

  • Inventory is the truth; live state wins over narrative docs.
  • Prefer homelab CLI and MCP over ad-hoc SSH.
  • Meaningful changes update docs in the same session.
  • Secrets are decrypted locally via per-client keys; never into docs/comments.
  • Tracked configs change by commit + push, not local edits.
  • Netbird is the preferred mesh path for new traffic.
  • Agents are terse (caveman.md), verify claims, and fix collateral drift when found.

Build status (30-day roadmap, started 2026-07-05)

  • Week 1: policy, ontology, service contract, archaeology, topology generator, this brief. Shipped.
  • Week 2: context cards, homelab service <name> …, change ledger, node relations, runbooks. Shipped.
  • Week 3: ops scheduler + state cache (homelab service <name> health is cache-first, --live forces a probe), drift detectors, signal engine (homelab signal …), decision classifier (homelab decide …), approval engine (homelab approval … — shared-HMAC grants; Matrix delivery is Hermes's existing @dtoro:avispero send path, not a new bot, see oikos/approve.py), daily brief + weekly report (oikos/report.py). Shipped, except: Prometheus is still planned (see plans/2026-07-05-oikos-prometheus-lxc.md) — trend signals (disk-full prediction, temp creep) wait on that LXC; the scheduler's disk check today is point-in-time only, and CPU/NVMe temperature isn't probed at all yet (no confirmed sensor path on hubris/strong). DNS-vs-inventory and generic tracked-config-cleanliness drift checks are also deferred (see oikos/drift.py docstring).
  • Week 4: Oikos Console v0 shipped — signals landing page, service grid + detail, node/blast-radius view, live Mermaid graph, drift view, approvals queue (approve/deny, destructive confirmation-phrase enforced), daily/weekly reports. Server-rendered FastAPI + Jinja2, no SPA build chain, tested end-to-end against live production data (see oikos/console/). Deploys as a third webhook on dtoro/Homelab-Docs (/opt/oikos-console, port :9831) — see oikos/console/deploy/README.md for the Caddy route and Gitea webhook registration this repo can't do for itself. Approval grants are now single-use (a second check_grant call for the same request fails even within the TTL) and already exact-bound to request id + entity + action. Not shipped as originally planned: per-agent age-key-signed request authentication — age has no signing primitive (it's an encryption-only keypair format), so "age-key-signed" wasn't buildable as stated. The real alternative (SSH-key signing via ssh-keygen -Y sign/-Y verify, using each host's already-provisioned SSH key) is real and buildable, but needs SSH public keys recorded in inventory first — not there today. Moved to the 60/90-day backlog. Authentik step-up re-auth on the approve/deny route is documented but needs a live Authentik instance to configure — also backlog. Docs pass done (this file, AGENTS.md, operations/commands.md); found and fixed two more stale references while at it (DNS section still pointed at destroyed LXC 124/dnsmasq instead of Technitium on 107, and a claudio-monitor reference that's been deprecated since 2026-06-04).

Real drift found while building Week 3 (unresolved, needs operator action)

The drift detectors surfaced genuine, currently-true findings on first run against production — recorded here rather than silently fixed, since each is a config_mutation/destructive-class decision:

  • republic-laptop has no age_pubkey: in inventory.yaml, but its real age key is granted on nearly every shared secret in .sops.yaml (age1vf8h7...) — the enrollment write-back to inventory never happened. Fix: homelab client add republic-laptop --finalize-pubkey age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6.
  • grimmory has an age_pubkey in inventory but is missing from secrets/hello.yaml's recipient list — incomplete enrollment the other direction. Fix: re-run homelab client add grimmory --finalize-pubkey <its key>.
  • pve_id 131 exists live on hubris (pct list) with no inventory entry — investigate before assuming it's a stale ID (see the Prometheus LXC plan doc above, which flags this explicitly).
  • Three lifecycle-pve-id-reuse info findings (100, 106, 107 each shared between an active host and an archaeology entry) — expected/benign ID reuse after destroy, no action needed.

60/90-day backlog

Derived from gaps observed while building the 30-day roadmap, not guesswork. Roughly ordered by what unblocks the most:

  • Fix the oikos-console deploy webhook's signature mismatch. Console is live on apps (105) via a manual deploy.sh run, but Gitea webhook 14's deliveries all 403 with a signature mismatch for a cause not yet found — the secret is confirmed synced correctly on both sides (rotated once already to rule out drift). Until fixed, git push doesn't auto-redeploy the console the way it does for homelab-mcp/ secrets-issuance; re-run deploy.sh on apps manually after changes. See oikos/console/deploy/README.md.
  • SSH-key-signed approval requests. Replaces the design note in Week 4: age keys can't sign (encryption-only format), so per-agent request authentication needs ssh-keygen -Y sign/-Y verify against each host's existing SSH key. Blocked on a schema gap: inventory doesn't record SSH public keys today, only ports/users. First step is populating that field on enrollment, then wiring oikos/approve.py to require and verify a signature over the request payload.
  • Authentik step-up re-auth on the Console's /approvals POST route — needs a live Authentik PromptStage/reauth flow scoped to that path; not configurable without a running instance to test against.
  • Prometheus provisioning (see plans/2026-07-05-oikos-prometheus-lxc.md) — unblocks trend signals (disk-full prediction, temp creep) and real sparklines in the Console; investigate the undocumented pve_id 131 on hubris first.
  • CPU/NVMe temperature probing in the scheduler — needs a confirmed sensor path on hubris and strong (lm-sensors vs vendor tool) before a real check can be written; guessing one risks a probe that silently never fires.
  • DNS-vs-inventory drift check — compare Technitium zone records against services.*.url/public_host; not implemented (oikos/drift.py has no Technitium API wiring yet).
  • Generic tracked-config-cleanliness drift check — today only caddy's /etc/caddy git-checkout path is hardcoded in oikos/drift.py; every other service with a config_repo needs its local checkout path recorded (a mutation_path-style field, same gap Week 1's service contract flagged but didn't backfill) before this generalizes.
  • Per-service policy overrides (oikos/policy.yaml service_overrides) — schema is ready (caddy/dns already use it); populate more as specific services turn out to need non-default risk classes.
  • Incident timeline generator — stitch ledger + signal history into a single narrative for investigations/ entries instead of writing them by hand.
  • Secret access audit — who-can-decrypt-what report from .sops.yaml + inventory age_pubkeys, extending what oikos/drift.py's SOPS check already partially does.
  • Restore drills — exercise backs-up-to (once populated) by actually restoring from a backup target on a schedule, not just checking freshness.
  • Multi-agent delegation model — more than one agent acting concurrently; needs the ledger's agent field to carry real identity (age pubkey, not just hostname) consistently, which it mostly does already but hasn't been stress-tested with concurrent writers.
  • Grafana — only if the Console's own Prometheus-backed sparklines turn out to be insufficient once Prometheus ships.
  • "Generalize later" extraction — the original decision was personal- first, generalize-later (see Week 1). Once patterns stabilize, extract a config-driven Oikos core with no hubris.network/hubris/strong hardcoding, so it's installable on a different homelab.