Problem: node and cross-cutting narratives lived at the repo root
(containers/, vms/, infrastructure/, host .md files), interleaved with the
machine-readable substrate.
Change:
- Move containers/ -> knowledge/wiki/containers/, vms/ -> knowledge/wiki/vms/,
infrastructure/ -> knowledge/wiki/infrastructure/, hosts/{hubris,strong}.md ->
knowledge/wiki/hosts/, infrastructure/references/ -> knowledge/sources/references/,
GLOSSARY.md -> knowledge/GLOSSARY.md.
- Add knowledge/{index.md,log.md,sources/index.md} scaffolding.
- Rewrite all relative links repo-wide via a path-resolving mapper (inbound +
outbound + between-moved-files), including .hermes/, runbooks, operations,
investigations, plans, README, AGENTS.
- Repoint inventory.yaml doc_page fields and regenerate hosts/*.yaml (which
embed doc_page); update oikos/gen-topology.py output path, candidate doc
paths, and footer links; update code-comment doc paths.
Substrate untouched in place: inventory.yaml, hosts/*.yaml (regenerated,
idempotent), oikos/ code, mcp/, secrets/, bin/.
Verification:
- Logical broken-link set identical to pre-move baseline (net 128 -> 127; the
topology regen fixed one, introduced none). Remaining are pre-existing refs
to destroyed/archived nodes, out of scope for this move.
- gen-topology.py --check exit 0 (in sync); cards carry knowledge/wiki/ doc paths.
- build_host_files.py idempotent; all inventory doc_page targets resolve.
- MCP contract verified: get_page/search_docs/get_changelog resolve moved pages.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
13 KiB
Oikos — the operating model
Oikos (Greek: household) is the agent operating system layered on this
repo. It is not new infrastructure: inventory.yaml is the kernel data
structure, the homelab CLI and MCP server are the syscall surface, and
this page defines the rules everything above them follows.
Read this after AGENTS.md. Machine-readable companions: oikos/ontology.yaml (systems model), oikos/policy.yaml (risk & approval).
The kernel loop: OODA
Every Oikos activity — scheduled probe, agent task, operator request — is one pass through Observe → Orient → Decide → Act:
- Observe — probes, drift detectors, and agent findings produce Signals (structured records, not loose messages): pending updates, high temperature, low disk, service down, cert expiry, stale backup, inventory drift.
- Orient — walk the ontology graph: what entity is affected, what depends on it (blast radius), its lifecycle state, whether a runbook matches, what the ledger says about past attempts.
- Decide — the classifier scores risk class × blast radius ×
confidence and routes:
- auto-act: within autonomy policy, high confidence, contained radius
- escalate: operator approval via Matrix (✅/❌ reaction) or the
Oikos Console's
/approvalspage (destructive actions additionally need a typed confirmation phrase either way) - queue: informational — console + reports The classifier can only lower autonomy relative to policy, never raise it. When in doubt, escalate.
- Act — execute through
homelabcommands or runbooks (never ad-hoc SSH), then verify with the action's verification command, write a ledger entry, resolve the Signal, and update docs in the same session.
Primitives
| Primitive | What it is | Lives in |
|---|---|---|
| Host / Service | topology entities | inventory.yaml (+ generated hosts/*.yaml) |
| Secret | SOPS+age encrypted value, per-client recipients | secrets/ + .sops.yaml |
| Runbook | executable workflow with risk class + verification | runbooks/ (Week 2) |
| Signal | something needing attention, with lifecycle | signals/ ledger (Week 3) |
| Change | one mutation: who, what, risk, approval, verification | ledger/ (Week 2) |
| Approval | short-TTL signed grant for a gated action | approval engine (Week 3) |
| Incident | investigation narrative | investigations/ |
| Plan | design doc for non-trivial work | plans/ |
| Agent | enrolled client identity = its age pubkey | inventory.yaml + .sops.yaml |
Risk classes (enforced, not advisory)
From oikos/policy.yaml:
- read_only — status, logs, docs, inventory. Unattended.
- reversible_low — restart, cache clear, sync pull. Unattended + ledger.
- config_mutation — tracked-config edits (commit+push, never local), deploys, upgrades, DNS/ingress changes. Operator approval.
- destructive — destroy, format, wipe, rotate, revoke. Approval + typed confirmation phrase.
Lifecycle gates modify these: provisioning nodes are freely mutable
(nothing depends on them); deprecated nodes accept no new dependents;
anything touching a destroyed node is drift.
The systems model
Eight domains — physical, compute, network, storage, software,
identity & access, operations, external — cover everything in the lab;
entities are connected by typed edges (hosts, provides, mounts,
stores-on, routes-to, can-decrypt, depends-on, backs-up-to, …)
defined in oikos/ontology.yaml. Rule of
completeness: if it can break, be changed, or hold data, it has an
entity and edges. Blast-radius questions ("what breaks if strong goes
down?") are graph walks, not doc archaeology.
Nodes move through an explicit lifecycle —
planned → provisioning → active → migrating → deprecated → destroyed —
stored as state: in inventory (absent = active). Destroyed nodes live in
the archaeology: section. Each transition is a runbook checklist;
deprecation completes only when inbound edges reach zero.
Generated views: infrastructure/topology.md
(Mermaid, regenerated from inventory) and the live, clickable version at
oikos.hubris.network/graph once the Console is deployed.
Conventions carried forward
- Inventory is the truth; live state wins over narrative docs.
- Prefer
homelabCLI and MCP over ad-hoc SSH. - Meaningful changes update docs in the same session.
- Secrets are decrypted locally via per-client keys; never into docs/comments.
- Tracked configs change by commit + push, not local edits.
- Netbird is the preferred mesh path for new traffic.
- Agents are terse (caveman.md), verify claims, and fix collateral drift when found.
Build status (30-day roadmap, started 2026-07-05)
- Week 1: policy, ontology, service contract, archaeology, topology generator, this brief. Shipped.
- Week 2: context cards,
homelab service <name> …, change ledger,node relations, runbooks. Shipped. - Week 3: ops scheduler + state cache (
homelab service <name> healthis cache-first,--liveforces a probe), drift detectors, signal engine (homelab signal …), decision classifier (homelab decide …), approval engine (homelab approval …— shared-HMAC grants; Matrix delivery is Hermes's existing@dtoro:avisperosend path, not a new bot, seeoikos/approve.py), daily brief + weekly report (oikos/report.py). Shipped, except: Prometheus is stillplanned(see plans/2026-07-05-oikos-prometheus-lxc.md) — trend signals (disk-full prediction, temp creep) wait on that LXC; the scheduler's disk check today is point-in-time only, and CPU/NVMe temperature isn't probed at all yet (no confirmed sensor path on hubris/strong). DNS-vs-inventory and generic tracked-config-cleanliness drift checks are also deferred (seeoikos/drift.pydocstring). - Week 4: Oikos Console v0 shipped — signals landing page, service
grid + detail, node/blast-radius view, live Mermaid graph, drift view,
approvals queue (approve/deny, destructive confirmation-phrase
enforced), daily/weekly reports. Server-rendered FastAPI + Jinja2, no
SPA build chain, tested end-to-end against live production data (see
oikos/console/). Deploys as a third webhook ondtoro/Homelab-Docs(/opt/oikos-console, port :9831) — see oikos/console/deploy/README.md for the Caddy route and Gitea webhook registration this repo can't do for itself. Approval grants are now single-use (a secondcheck_grantcall for the same request fails even within the TTL) and already exact-bound to request id + entity + action. Not shipped as originally planned: per-agent age-key-signed request authentication — age has no signing primitive (it's an encryption-only keypair format), so "age-key-signed" wasn't buildable as stated. The real alternative (SSH-key signing viassh-keygen -Y sign/-Y verify, using each host's already-provisioned SSH key) is real and buildable, but needs SSH public keys recorded in inventory first — not there today. Moved to the 60/90-day backlog. Authentik step-up re-auth on the approve/deny route is documented but needs a live Authentik instance to configure — also backlog. Docs pass done (this file, AGENTS.md, operations/commands.md); found and fixed two more stale references while at it (DNS section still pointed at destroyed LXC 124/dnsmasq instead of Technitium on 107, and aclaudio-monitorreference that's been deprecated since 2026-06-04).
Real drift found while building Week 3 (unresolved, needs operator action)
The drift detectors surfaced genuine, currently-true findings on first
run against production — recorded here rather than silently fixed, since
each is a config_mutation/destructive-class decision:
republic-laptophas noage_pubkey:ininventory.yaml, but its real age key is granted on nearly every shared secret in.sops.yaml(age1vf8h7...) — the enrollment write-back to inventory never happened. Fix:homelab client add republic-laptop --finalize-pubkey age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6.grimmoryhas anage_pubkeyin inventory but is missing fromsecrets/hello.yaml's recipient list — incomplete enrollment the other direction. Fix: re-runhomelab client add grimmory --finalize-pubkey <its key>.pve_id 131exists live on hubris (pct list) with no inventory entry — investigate before assuming it's a stale ID (see the Prometheus LXC plan doc above, which flags this explicitly).- Three
lifecycle-pve-id-reuseinfo findings (100, 106, 107 each shared between an active host and an archaeology entry) — expected/benign ID reuse after destroy, no action needed.
60/90-day backlog
Derived from gaps observed while building the 30-day roadmap, not guesswork. Roughly ordered by what unblocks the most:
- Fix the oikos-console deploy webhook's signature mismatch. Console
is live on apps (105) via a manual
deploy.shrun, but Gitea webhook 14's deliveries all 403 with a signature mismatch for a cause not yet found — the secret is confirmed synced correctly on both sides (rotated once already to rule out drift). Until fixed,git pushdoesn't auto-redeploy the console the way it does for homelab-mcp/ secrets-issuance; re-rundeploy.shon apps manually after changes. See oikos/console/deploy/README.md. - SSH-key-signed approval requests. Replaces the design note in
Week 4: age keys can't sign (encryption-only format), so per-agent
request authentication needs
ssh-keygen -Y sign/-Y verifyagainst each host's existing SSH key. Blocked on a schema gap: inventory doesn't record SSH public keys today, only ports/users. First step is populating that field on enrollment, then wiringoikos/approve.pyto require and verify a signature over the request payload. - Authentik step-up re-auth on the Console's
/approvalsPOST route — needs a live AuthentikPromptStage/reauth flow scoped to that path; not configurable without a running instance to test against. - Prometheus provisioning (see
plans/2026-07-05-oikos-prometheus-lxc.md)
— unblocks trend signals (disk-full prediction, temp creep) and real
sparklines in the Console; investigate the undocumented
pve_id 131on hubris first. - CPU/NVMe temperature probing in the scheduler — needs a confirmed sensor path on hubris and strong (lm-sensors vs vendor tool) before a real check can be written; guessing one risks a probe that silently never fires.
- DNS-vs-inventory drift check — compare Technitium zone records
against
services.*.url/public_host; not implemented (oikos/drift.pyhas no Technitium API wiring yet). - Generic tracked-config-cleanliness drift check — today only caddy's
/etc/caddygit-checkout path is hardcoded inoikos/drift.py; every other service with aconfig_reponeeds its local checkout path recorded (amutation_path-style field, same gap Week 1's service contract flagged but didn't backfill) before this generalizes. - Per-service policy overrides (
oikos/policy.yamlservice_overrides) — schema is ready (caddy/dns already use it); populate more as specific services turn out to need non-default risk classes. - Incident timeline generator — stitch ledger + signal history into
a single narrative for
investigations/entries instead of writing them by hand. - Secret access audit — who-can-decrypt-what report from
.sops.yaml+ inventoryage_pubkeys, extending whatoikos/drift.py's SOPS check already partially does. - Restore drills — exercise
backs-up-to(once populated) by actually restoring from a backup target on a schedule, not just checking freshness. - Multi-agent delegation model — more than one agent acting
concurrently; needs the ledger's
agentfield to carry real identity (age pubkey, not just hostname) consistently, which it mostly does already but hasn't been stress-tested with concurrent writers. - Grafana — only if the Console's own Prometheus-backed sparklines turn out to be insufficient once Prometheus ships.
- "Generalize later" extraction — the original decision was personal-
first, generalize-later (see Week 1). Once patterns stabilize, extract
a config-driven Oikos core with no
hubris.network/hubris/stronghardcoding, so it's installable on a different homelab.