# Plan: Oikos — Docker-based agentic homelab OS on mac-mini **Status:** Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first design with systems modeling, DB-native config (inventory/ontology/policy as graph metadata, not YAML files), and a feedback loop where the agent learns from execution. ## Vision Convert this repo into a **Docker-based agentic homelab OS** written in **Go**. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors: - **Operator** (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix. - **Agent** — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and **learns from every action** to improve over time. All OS services run in Docker containers on mac-mini. The OS is deployed by a git push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path to multi-node later. ## Decisions (from operator Q&A, 2026-07-06) | Question | Decision | |---|---| | Repo structure | One repo, reorganize internally. | | Language | **Go** — compiled, type-safe, small containers, goroutines for concurrent probes. | | Hermes runtime | Runs inside Docker as part of the OS stack (gateway mode). | | Agent → homelab access | Hybrid — mounted SSH keys now, actuator gateway built incrementally. | | Data storage | PostgreSQL. Inventory/ontology/policy become DB-native graph metadata. | | Config (inventory/ontology/policy) | **DB-native** — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend. | | Knowledge/context | Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data. | | Operator interface | Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks. | | Host | mac-mini for now, designed to scale later. | | Deploy | Git push → Gitea webhook → Docker rebuild + restart. | | `bin/homelab` CLI | Replaced by an API. CLI becomes a thin Go client that calls the OS API. | | Secrets | Migrate from SOPS+age to Infisical. | | Matrix | Keep for now, abstract the notification layer for future channels. | | MCP server | Merged into the unified API — one Go service, REST + MCP interfaces. | | Agent type | Hermes core + custom homelab skills. | | Feedback loop | **Agent learns from execution** — outcomes feed back as patterns and skills that inform future decisions. | | Ontology | **Developed first**, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles. | | apps/105 | Keep running as fallback until Docker OS is proven. | | mac-mini cleanup | Start fresh in a new directory, clean up old artifacts later. | ## Ontology — the systems model The ontology is the foundational layer of the OS. It defines what exists, how things connect, how they change over time, and how the OS learns about them. It is stored **in the database** as metadata (entity types, relationship types, lifecycle definitions). YAML seed files bootstrap it on first deploy; after that, the DB is authoritative and editable via API. ### Design principles (systems modeling) 1. **Three layers** — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's own behavior + learning). Dependencies flow upward: Cognition depends on Governance depends on Infrastructure. 2. **Everything is an entity** — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, changes, skills) are first-class entities, not second-class records. 3. **Typed relationships with cardinality** — edges carry semantics. `hosts` is one-to-many; `depends-on` is many-to-many; `documents` is one-to-one. The graph is queryable for blast radius, dependency chains, and knowledge lookup. 4. **Lifecycles are state machines** — every entity type has a lifecycle. Infrastructure entities move through `planned → active → destroyed`. Operational entities have their own lifecycles (signals, approvals, patterns, skills). Transitions can require preconditions. 5. **Policy is attached to the ontology** — risk classes and approval rules link to entity types and actions. The policy IS part of the model, not a separate file. 6. **Learning is modeled** — executions produce outcomes, outcomes accumulate into patterns, patterns refine skills, skills inform future decisions. This is an explicit, queryable part of the graph. ### Layer map ```mermaid graph TB subgraph cognition["Layer 3 — Cognition (the OS's behavior + learning)"] direction LR OBS["Observation\nsignal, state-snapshot"] DEC["Decision\nclassification, risk-assessment"] ACT["Action\nchange, execution, verification"] GOV["Governance\napproval-request, approval-decision"] KNOW["Knowledge\ndocument, runbook, lesson"] LEARN["Learning\npattern, skill, feedback"] end subgraph governance["Layer 2 — Governance (who controls what)"] direction LR IDENT["Identity\nperson, agent, identity-provider"] SEC["Secrets\nsecret, key, access-grant"] POL["Policy\nrisk-class, approval-rule, autonomy-setting"] end subgraph infra["Layer 1 — Infrastructure (the managed world)"] direction LR PHYS["Physical\nsite, machine, ups, sensor"] COMP["Compute\nproxmox-host, lxc, vm,\nworkstation, container"] NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"] STOR["Storage\nstorage-pool, volume,\nmount, backup-target"] SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"] end cognition -- "observes, acts on, learns about" --> infra governance -- "governs access to" --> infra governance -- "constrains" --> cognition cognition -- "creates + refines" --> governance ``` ### Entity relationship diagram — core entities and typed edges ```mermaid erDiagram MACHINE ||--o{ PROXMOX_HOST : "is-a" PROXMOX_HOST ||--o{ LXC : hosts PROXMOX_HOST ||--o{ VM : hosts PROXMOX_HOST ||--o{ WORKSTATION : hosts LXC ||--o{ SERVICE : provides VM ||--o{ SERVICE : provides WORKSTATION ||--o{ SERVICE : provides LXC ||--o{ MOUNT : has MOUNT }o--|| STORAGE_POOL : "stores-on" SERVICE ||--o{ INGRESS_ROUTE : "exposed-by" INGRESS_ROUTE }o--|| CERTIFICATE : "secured-by" INGRESS_ROUTE }o--|| IDENTITY_PROVIDER : "secured-by" SERVICE ||--o{ SERVICE : "depends-on" CONFIG_REPO ||--|| SERVICE : "configured-by" DEPLOY_PIPELINE ||--|| SERVICE : "deploys-to" SERVICE ||--o{ SIGNAL : "monitored-by" SIGNAL ||--o{ EXECUTION : triggers EXECUTION ||--|| FEEDBACK : produces FEEDBACK }o--|| PATTERN : "contributes-to" PATTERN }o--|| SKILL : "informs" SKILL ||--o{ CLASSIFICATION : "guides" CLASSIFICATION ||--|| EXECUTION : "precedes" EXECUTION ||--o| APPROVAL : "requires" PERSON ||--o{ APPROVAL : "decides" DOCUMENT ||--o{ SERVICE : "describes" RUNBOOK ||--o{ SERVICE : "procedure-for" AGENT ||--o{ EXECUTION : "performs" PERSON ||--o{ AGENT : "owns" ``` ### Infrastructure lifecycle ```mermaid stateDiagram-v2 [*] --> planned : operator creates entity planned --> provisioning : IP reserved, storage chosen, doc stub provisioning --> active : mesh joined, health-check answering, doc complete active --> migrating : preflight + backup verified migrating --> active : post-verify, caddy checked, mounts checked active --> deprecated : replacement live or role retired deprecated --> destroyed : backups verified, secrets revoked, ingress removed destroyed --> [*] : archaeology entry recorded note right of deprecated Complete only when zero inbound edges remain (no depends-on, no routes-to) end note ``` ### Signal lifecycle ```mermaid stateDiagram-v2 [*] --> raised : scheduler probe or agent finding raised --> acknowledged : agent or operator sees it raised --> muted : operator suppresses (TTL) acknowledged --> acting : actuator starts execution acting --> resolved : action succeeded, verification passed acting --> raised : action failed, re-escalated raised --> resolved : condition cleared (auto-resolve) muted --> raised : TTL expired resolved --> [*] ``` ### Execution + learning lifecycle ```mermaid stateDiagram-v2 [*] --> proposed : signal + recommended_action proposed --> approved : operator approves (if gated) proposed --> auto_approved : risk class allows auto-act approved --> executing : actuator runs auto_approved --> executing : actuator runs executing --> verified : verification command passed executing --> failed : execution or verification failed executing --> timed_out : exceeded duration limit failed --> rolled_back : rollback procedure executed verified --> [*] : feedback recorded, pattern updated failed --> [*] : feedback recorded, pattern updated rolled_back --> [*] : feedback recorded, pattern updated timed_out --> [*] : feedback recorded, pattern updated ``` ### Feedback / learning model — the cognition loop ```mermaid flowchart TB subgraph observe["Observe"] SIG["Signal raised\n(service down, disk full, drift)"] end subgraph decide["Decide"] CLASS["Classification\nrisk × blast radius × confidence"] SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"] CLASS --> SKILL_LOOKUP end subgraph act["Act"] EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"] VERIFY["Verification\ncheck if action succeeded"] EXEC --> VERIFY end subgraph learn["Learn (feedback loop)"] OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"] FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"] PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"] SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"] OUTCOME --> FEEDBACK FEEDBACK --> PATTERN PATTERN --> SKILL_REFINE end SIG --> CLASS SKILL_LOOKUP --> EXEC VERIFY --> OUTCOME SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP subgraph escalation["Escalation path"] APPROVAL["Approval request\n→ Matrix ✅/❌"] end CLASS -- "needs approval" --> APPROVAL APPROVAL -- "approved" --> EXEC APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"] ``` ### Policy model — how risk attaches to entities ```mermaid flowchart LR subgraph ontology["Ontology (DB metadata)"] ET["entity_types\nhost, service, lxc, vm, ..."] RT["relationship_types\nhosts, provides, depends-on, ..."] LD["lifecycle_defs\nstates + transitions"] end subgraph policy["Policy (DB records)"] RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"] RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"] AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"] end subgraph instances["Instances (DB data)"] ENT["entities\nhost:hubris, service:caddy, ..."] REL["relationships\nhubris hosts lxc:apps"] end ET --> RULES RC --> RULES RULES --> ENT AUTO --> ENT ET --> ENT RT --> REL LD --> ENT ``` ## Target architecture ### Container stack on mac-mini ```mermaid graph TB subgraph mac-mini["mac-mini — Docker host, always-on"] subgraph services["Docker Compose — Go binaries"] PG["PostgreSQL 16\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata"] INF["Infisical\n(secrets manager)"] HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"] API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events\n\nPolicy enforcement + risk\nclassification + ledger"] SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• SSH to hubris/strong"] NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"] end DEPLOY["Gitea webhook →\ndocker compose build + up -d"] end HERMES -- MCP --> API API --> PG SCHED --> PG SCHED -- SSH --> HUBRIS SCHED -- SSH --> STRONG API -- "escalate" --> NOTIFIER NOTIFIER -- "alerts + approvals" --> MATRIX DEPLOY -- "rebuild" --> services subgraph external["External"] CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"] APPS["apps/105 (fallback)"] HUBRIS["hubris (PVE)"] STRONG["strong (PVE)"] GITEA["Gitea (LXC 104)"] MATRIX["Matrix (LXC 118)"] WS["Any workstation\nHermes remote → gateway"] end CADDY -- reverse_proxy --> API GITEA -- webhook --> DEPLOY WS -- "Hermes gateway" --> HERMES ``` ### OODA loop — with learning feedback ```mermaid flowchart LR OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"] ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"] DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"] DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"] APPROVE -- "approved" --> ACT ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"] LEARN -. "improves confidence" .-> DECIDE LEARN --> OBSERVE ``` ### Deploy flow ```mermaid flowchart LR DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"] GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"] MACMINI -- "git pull" --> REPO["repo clone"] MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"] STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"] ``` ## DB-native configuration The three YAML files — `inventory.yaml`, `ontology.yaml`, `policy.yaml` — become **seed manifests**. They bootstrap the DB on first deploy. After that, the DB is the runtime source of truth, editable via the API. A future frontend can edit all three directly. ### How it works ```mermaid flowchart TB subgraph seeds["Seed manifests (git-tracked, YAML)"] ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"] INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"] POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"] end subgraph db["PostgreSQL (runtime source of truth)"] META["entity_types table\nrelationship_types table\nlifecycle_defs table"] INST["entities table\nrelationships table"] POLDB["policies table\nrisk_classes table\nautonomy_settings table"] end INGEST["Seed ingest (on deploy)\nidempotent upsert"] ONTO_YAML --> INGEST --> META INV_YAML --> INGEST --> INST POL_YAML --> INGEST --> POLDB API_EDIT["API edits\n(POST/PUT/PATCH)"] API_EDIT --> META API_EDIT --> INST API_EDIT --> POLDB EXPORT["Export to YAML\n(for DR / version control)"] META --> EXPORT INST --> EXPORT POLDB --> EXPORT ``` ### Why DB-native - **Querying** — the agent can ask "what services depend on authentik?" as a graph query, not a YAML parse. Blast-radius walks are SQL, not file reads. - **Mutation** — adding a service, updating a lifecycle state, changing a policy rule are DB transactions with audit trail, not file edits + git commits. - **Consistency** — the ontology, inventory, and policy are always in sync (same DB, same transaction). No drift between what the YAML says and what the runtime sees. - **Future frontend** — a UI can edit entities, relationships, and policies directly via the API. No need to generate/edit YAML files. - **Version control** — seed YAML files are still git-tracked for bootstrap and DR. The API can export the current DB state back to YAML for commit. ## Repo layout (Go project) ``` / # repo root ├── docker-compose.yml # the OS stack definition ├── Makefile # build, test, deploy targets ├── go.mod # Go module definition ├── go.sum ├── cmd/ # binary entrypoints (one per service) │ ├── api/ # Oikos API server │ │ └── main.go │ ├── scheduler/ # Observe + Act loop │ │ └── main.go │ └── notifier/ # Notification service │ └── main.go ├── internal/ # private packages (not importable) │ ├── db/ # database layer │ │ ├── queries/ # sqlc SQL queries │ │ ├── models.go # generated Go types │ │ └── db.go # connection pool, migrations │ ├── ontology/ # ontology types + meta-schema │ │ ├── types.go # EntityType, RelationshipType, LifecycleDef │ │ ├── graph.go # graph traversal (blast radius, dependencies) │ │ └── ingest.go # YAML seed → DB ingest │ ├── api/ # HTTP + MCP server │ │ ├── server.go # Gin app setup │ │ ├── routes/ # REST handlers │ │ │ ├── hosts.go │ │ │ ├── services.go │ │ │ ├── signals.go │ │ │ ├── approvals.go │ │ │ ├── exec.go │ │ │ └── knowledge.go │ │ └── mcp.go # MCP protocol adapter (JSON-RPC over SSE) │ ├── policy/ # risk classification + approval │ │ ├── classify.go # risk × blast × confidence │ │ ├── approve.go # approval request + grant lifecycle │ │ └── autonomy.go # kill-switch, never-auto-act list │ ├── scheduler/ # Observe stage │ │ ├── probe.go # HTTP health, disk, drift │ │ └── signal.go # raise/resolve signals in DB │ ├── actuator/ # Act stage │ │ ├── act.go # read signals, classify, execute or escalate │ │ ├── execute.go # SSH execution + verification │ │ └── guard.go # loop-guard, retry caps │ ├── learning/ # feedback loop (the learning model) │ │ ├── feedback.go # record outcome + lesson from execution │ │ ├── pattern.go # extract/validate patterns from feedback │ │ └── skill.go # create/refine skills from patterns │ ├── notifier/ # notification abstraction │ │ ├── notifier.go # interface │ │ └── matrix.go # Matrix implementation │ └── config/ # config loading (env, files) │ └── config.go ├── migrations/ # SQL migrations (golang-migrate format) │ ├── 001_ontology.up.sql # meta-schema (entity_types, relationship_types, ...) │ ├── 001_ontology.down.sql │ ├── 002_instances.up.sql # entities, relationships │ ├── 003_operations.up.sql # signals, approvals, executions, feedback │ ├── 004_learning.up.sql # patterns, skills │ └── 005_policy.up.sql # policies, risk_classes, autonomy ├── seeds/ # YAML seed manifests (bootstrap + DR) │ ├── ontology.yaml # entity types, relationship types, lifecycles │ ├── inventory.yaml # entity instances (hosts, services, etc.) │ └── policy.yaml # risk classes, approval rules, autonomy ├── docs/ # narrative docs (ingested into knowledge graph) │ ├── containers/ │ ├── hosts/ │ ├── infrastructure/ │ └── investigations/ ├── hermes/ # Hermes agent config + skills │ ├── config.yaml │ ├── SOUL.md │ └── skills/ │ └── homelab-ops/ │ └── SKILL.md ├── compose/ # Docker build contexts │ ├── api/Dockerfile │ ├── scheduler/Dockerfile │ ├── hermes/Dockerfile │ └── postgres/init.sql └── scripts/ # utility scripts ├── migrate-sops.sh # one-time SOPS → Infisical migration └── import-legacy.sh # import existing signals/ledger JSONL ``` ## Database schema The schema is the ontology made concrete. Five migration groups, each adding a layer. ### Migration 1: Ontology meta-schema ```sql -- The meta-graph: defines what entity types and relationship types can exist. -- This IS the ontology, stored in the DB, editable via API. CREATE TABLE entity_types ( name TEXT PRIMARY KEY, -- 'host', 'service', 'signal', 'pattern' domain TEXT NOT NULL, -- 'physical', 'compute', 'network', ... layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition' description TEXT, lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle) created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE TABLE relationship_types ( name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on' inverse TEXT, -- 'runs-on', 'provided-by' source_type TEXT REFERENCES entity_types(name), target_type TEXT REFERENCES entity_types(name), cardinality TEXT NOT NULL, -- 'one-to-one', 'one-to-many', 'many-to-many' description TEXT, created_at TIMESTAMPTZ DEFAULT now() ); CREATE TABLE lifecycle_defs ( id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', 'execution', ... states TEXT[] NOT NULL, -- ordered states default_state TEXT NOT NULL, transitions JSONB NOT NULL, -- {"from": {"to": {"requires": [...]}}} created_at TIMESTAMPTZ DEFAULT now() ); ``` ### Migration 2: Entity instances (the inventory graph) ```sql -- Entity instances — the actual hosts, services, signals, patterns, etc. -- This replaces inventory.yaml as the runtime source of truth. CREATE TABLE entities ( id TEXT PRIMARY KEY, -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001' type TEXT NOT NULL REFERENCES entity_types(name), name TEXT NOT NULL, -- 'hubris', 'caddy', 'disk-threshold' state TEXT, -- lifecycle state (e.g. 'active', 'raised') attributes JSONB NOT NULL DEFAULT '{}', -- type-specific data (IP, mesh addr, port, ...) parent_id TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host) created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE INDEX idx_entities_type ON entities(type); CREATE INDEX idx_entities_state ON entities(state); CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes); -- Relationship instances — the typed edges of the graph CREATE TABLE relationships ( source_id TEXT NOT NULL REFERENCES entities(id), target_id TEXT NOT NULL REFERENCES entities(id), type TEXT NOT NULL REFERENCES relationship_types(name), attributes JSONB, created_at TIMESTAMPTZ DEFAULT now(), PRIMARY KEY (source_id, target_id, type) ); CREATE INDEX idx_rel_source ON relationships(source_id); CREATE INDEX idx_rel_target ON relationships(target_id); CREATE INDEX idx_rel_type ON relationships(type); -- Recursive graph traversal function (blast radius, dependency chains) CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3) RETURNS TABLE(entity_id TEXT, depth INT) AS $$ WITH RECURSIVE walk AS ( SELECT start_id::TEXT AS entity_id, 0::INT AS depth UNION SELECT r.target_id::TEXT, w.depth + 1 FROM relationships r JOIN walk w ON r.source_id = w.entity_id WHERE w.depth < max_depth ) SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id; $$ LANGUAGE sql STABLE; ``` ### Migration 3: Operations (signals, approvals, ledger, state) ```sql -- Signals — now entities in the graph, with a dedicated table for indexed querying -- (the entity row is the canonical record; this table is a fast lookup) CREATE TABLE signals ( entity_id TEXT PRIMARY KEY REFERENCES entities(id), kind TEXT NOT NULL, severity TEXT NOT NULL, -- info, warning, critical target_entity_id TEXT REFERENCES entities(id), -- the infrastructure entity this is about evidence TEXT, likely_cause TEXT, recommended_action JSONB, verification TEXT, state TEXT NOT NULL DEFAULT 'raised', mute_until TIMESTAMPTZ, created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE INDEX idx_signals_state ON signals(state); CREATE INDEX idx_signals_target ON signals(target_entity_id); CREATE INDEX idx_signals_severity ON signals(severity); -- Approvals CREATE TABLE approvals ( id TEXT PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), -- entity to act on action TEXT NOT NULL, risk_class TEXT NOT NULL, status TEXT NOT NULL DEFAULT 'pending', -- pending, approved, denied, expired ttl INTERVAL NOT NULL DEFAULT '1 hour', decided_at TIMESTAMPTZ, decided_by TEXT REFERENCES entities(id), -- person entity confirmation_phrase TEXT ); -- Change ledger (high-level record, links to execution for detail) CREATE TABLE ledger_entries ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), action TEXT NOT NULL, risk_class TEXT NOT NULL, result TEXT, -- ok, failed, escalated approval_id TEXT REFERENCES approvals(id), execution_id INTEGER, -- FK to executions (migration 4) agent_id TEXT REFERENCES entities(id), notes TEXT ); -- State snapshots (replaces oikos/state.json) CREATE TABLE state_snapshots ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), health TEXT, -- healthy, degraded, down, unknown data JSONB ); ``` ### Migration 4: Learning model (executions, feedback, patterns, skills) ```sql -- Executions — detailed record of each action the OS performs CREATE TABLE executions ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), signal_entity_id TEXT REFERENCES entities(id), -- signal that triggered this target_entity_id TEXT REFERENCES entities(id), -- entity acted upon action TEXT NOT NULL, risk_class TEXT NOT NULL, approval_id TEXT REFERENCES approvals(id), agent_id TEXT REFERENCES entities(id), -- who/what executed skill_id TEXT REFERENCES entities(id), -- skill used (if any) status TEXT NOT NULL DEFAULT 'queued', -- queued, running, completed, failed, timed-out result JSONB, -- detailed result data duration_ms INTEGER, verified BOOLEAN DEFAULT false, started_at TIMESTAMPTZ, completed_at TIMESTAMPTZ ); CREATE INDEX idx_exec_target ON executions(target_entity_id); CREATE INDEX idx_exec_status ON executions(status); CREATE INDEX idx_exec_action ON executions(action); -- Feedback — what was learned from an execution CREATE TABLE feedback ( id SERIAL PRIMARY KEY, execution_id INTEGER REFERENCES executions(id), ts TIMESTAMPTZ DEFAULT now(), outcome TEXT NOT NULL, -- success, failure, partial, unexpected observation TEXT, -- what happened vs what was expected lesson TEXT, -- extractable lesson unexpected_side_effects TEXT[], tags TEXT[] ); CREATE INDEX idx_feedback_execution ON feedback(execution_id); CREATE INDEX idx_feedback_outcome ON feedback(outcome); -- Patterns — generalized rules extracted from accumulated feedback CREATE TABLE patterns ( id TEXT PRIMARY KEY, -- 'pat-2026-07-06-001' ts TIMESTAMPTZ DEFAULT now(), entity_type TEXT REFERENCES entity_types(name), -- applies to this type action TEXT NOT NULL, -- 'restart', 'deploy', etc. pattern TEXT NOT NULL, -- 'service X recovers within 30s after restart' confidence REAL DEFAULT 0.5, -- 0.0 to 1.0 evidence_count INTEGER DEFAULT 1, -- how many executions support this success_count INTEGER DEFAULT 0, failure_count INTEGER DEFAULT 0, status TEXT DEFAULT 'hypothesized', -- hypothesized, validated, active, deprecated last_validated_at TIMESTAMPTZ ); CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action); CREATE INDEX idx_patterns_status ON patterns(status); -- Skills — codified procedures refined through feedback CREATE TABLE skills ( id TEXT PRIMARY KEY, -- 'skill-restart-service', 'skill-deploy-lxc' name TEXT NOT NULL, ts TIMESTAMPTZ DEFAULT now(), procedure TEXT NOT NULL, -- the codified steps (markdown or structured) applies_to TEXT REFERENCES entity_types(name), pattern_ids TEXT[], -- patterns that inform this skill status TEXT DEFAULT 'drafted', -- drafted, tested, active, refined, deprecated version INTEGER DEFAULT 1, success_rate REAL, -- rolling success rate last_used_at TIMESTAMPTZ ); CREATE INDEX idx_skills_type ON skills(applies_to); CREATE INDEX idx_skills_status ON skills(status); ``` ### Migration 5: Policy (DB-native risk + approval rules) ```sql -- Risk classes — the four-level safety model CREATE TABLE risk_classes ( name TEXT PRIMARY KEY, -- 'read_only', 'reversible_low', etc. description TEXT, approval_required TEXT NOT NULL DEFAULT 'none', -- none, operator, operator_confirmed ledger BOOLEAN DEFAULT false, autonomy_allowed BOOLEAN DEFAULT false -- can agent auto-act at this risk level? ); -- Approval rules — entity_type + action → risk_class + requirements CREATE TABLE approval_rules ( id SERIAL PRIMARY KEY, entity_type TEXT REFERENCES entity_types(name), -- applies to this entity type action TEXT NOT NULL, -- 'restart', 'deploy', 'destroy' risk_class TEXT NOT NULL REFERENCES risk_classes(name), autonomy_level TEXT NOT NULL DEFAULT 'auto', -- 'auto', 'escalate', 'never' scope_entity TEXT REFERENCES entities(id), -- optional: specific entity only created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now(), UNIQUE(entity_type, action) ); -- Autonomy settings — global kill-switch + per-entity overrides CREATE TABLE autonomy_settings ( key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.caddy' value TEXT NOT NULL, -- 'off', 'reversible_low', 'true', 'false' updated_at TIMESTAMPTZ DEFAULT now() ); ``` ## Workstreams ### 1. Ontology definition + seed manifests (`seeds/`, `internal/ontology/`) **Before any code**, finalize the ontology. The current `ontology.yaml` has 8 domains and 14 relationship types. The new ontology adds: - **Layer 3 entities** (cognition): signal, execution, feedback, pattern, skill, approval, classification, document, runbook - **New relationships**: `triggers`, `produces`, `contributes-to`, `informs`, `guides`, `precedes`, `performs`, `procedure-for`, `learned-from` - **Lifecycle definitions** for operational entities (signals, executions, patterns, skills) — not just infrastructure **Deliverables:** - `seeds/ontology.yaml` — entity types, relationship types, lifecycle definitions (seeded into `entity_types`, `relationship_types`, `lifecycle_defs` on deploy) - `seeds/inventory.yaml` — adapted from current `inventory.yaml` (seeded into `entities` + `relationships`) - `seeds/policy.yaml` — adapted from current `policy.yaml` (seeded into `risk_classes`, `approval_rules`, `autonomy_settings`) - `internal/ontology/ingest.go` — idempotent seed → DB ingest ### 2. Database layer + migrations (`migrations/`, `internal/db/`) - Write the 5 migrations above - Set up sqlc for type-safe Go database access - Connection pool, migration runner - Graph traversal queries (blast radius, dependency chains, knowledge lookup) - Import script for existing `signals/*.jsonl`, `ledger/*.jsonl` → DB ### 3. Unified API server — Go (`cmd/api/`, `internal/api/`) One Go binary (Gin web framework) exposing REST + MCP from the same codebase. **MCP interface** (`internal/api/mcp.go`): - JSON-RPC over SSE, compatible with Hermes MCP client - Tools: `get_host`, `list_services`, `search_knowledge`, `get_entity`, `get_relations`, `get_blast_radius`, `get_signal_history`, `get_ledger`, `get_patterns`, `get_skills`, `get_state_snapshot` - All read from PostgreSQL **REST interface** (`internal/api/routes/`): - `GET /api/v1/entities` — list entities (filter by type, state, domain) - `GET /api/v1/entities/{id}` — entity detail + relationships - `POST /api/v1/entities` — create entity (creates inventory entry) - `PATCH /api/v1/entities/{id}` — update entity (state transition, attributes) - `GET /api/v1/signals` — list signals (filter by state, severity, entity) - `POST /api/v1/signals/{id}/ack` — acknowledge - `POST /api/v1/signals/{id}/resolve` — resolve - `GET /api/v1/approvals` — pending approvals - `POST /api/v1/approvals/{id}/decide` — approve/deny (Authentik-gated) - `POST /api/v1/exec` — gated execution (classify → check approval → execute → feedback) - `GET /api/v1/patterns` — list patterns (filter by entity_type, action, status) - `GET /api/v1/skills` — list skills (filter by applies_to, status) - `GET /api/v1/knowledge/{entity_id}` — knowledge graph query - `GET /api/v1/knowledge/search?q=...` — search knowledge graph - `GET /api/v1/ontology` — list entity types, relationship types, lifecycles - `POST /api/v1/ontology/entity_types` — create entity type (extend the schema) - `WS /api/v1/events` — real-time stream (signals, approvals, executions, feedback) **Policy enforcement** (`internal/policy/classify.go`): - Every mutating endpoint classifies the action via the policy DB - Risk class → approval check → autonomy check - All mutations write to the ledger automatically **Auth:** - MCP interface: no auth (internal, container-to-container) - REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints - Internal: shared secret (Docker network) ### 4. Scheduler + Actuator — Go (`cmd/scheduler/`, `internal/scheduler/`, `internal/actuator/`) **Scheduler (Observe)** — Go service with goroutines for concurrent probes: - HTTP health probes (concurrent, with timeouts) - Disk usage probes (SSH to hubris/strong) - Drift detection (inventory vs live state) - Writes signals + state snapshots to DB - Runs on a 10-min ticker **Actuator (Act)** — Go service, the control loop: - Reads open signals with `recommended_action` - For each: classify via `internal/policy/classify.go` - **auto-act**: look up skill for (entity_type, action) → follow procedure → execute via SSH → verify → record execution → generate feedback → update patterns - **escalate**: create approval request → notify via Matrix → acknowledge signal - Loop-guard: check execution history per (entity, action) to cap auto-retries (`SELECT ... FOR UPDATE SKIP LOCKED` for concurrency safety) - Autonomy kill-switch: check `autonomy_settings` table ### 5. Learning engine — Go (`internal/learning/`) The feedback loop that makes the agent improve over time. **Feedback recording** (`internal/learning/feedback.go`): - After every execution, evaluate the outcome: - Did the verification command pass? → success - Did it fail? → failure - Did it partially work? → partial - Did something unexpected happen? → unexpected - Record a feedback entry with: outcome, observation (what happened vs expected), lesson (extractable insight), unexpected_side_effects **Pattern extraction** (`internal/learning/pattern.go`): - Periodically scan accumulated feedback for (entity_type, action) pairs - When N+ executions share a similar outcome, extract a pattern: - "Restarting service:X typically takes 15s and succeeds" - "Deploying to LXC:Y via webhook has 30% failure rate, retry helps" - Patterns start as `hypothesized`, move to `validated` after enough evidence, then `active` (used by the decision classifier) - Confidence score = success_count / evidence_count, adjusted by recency **Skill management** (`internal/learning/skill.go`): - When a pattern reaches `active` status with confidence > 0.7, create or refine a skill for that (entity_type, action) pair - Skills codify the best-known procedure (what steps to take, what to verify, expected duration, known failure modes) - Skills are versioned — each refinement increments the version - The actuator looks up skills before executing: if a skill exists, follow it; if not, use the default procedure and generate feedback for future pattern extraction **How the classifier uses learning** (`internal/policy/classify.go`): - Confidence scoring now checks patterns + skills, not just raw ledger history: - If a pattern exists for (entity_type, action) with high confidence → boost auto-act confidence - If patterns show frequent failures → lower confidence, escalate - If a skill exists → higher confidence (proven procedure available) - This is the closed loop: **execution → feedback → pattern → skill → classification → execution** (better informed each time) ### 6. Knowledge graph ingestion (`internal/ontology/ingest.go`) After the ontology is defined and the DB schema is in place: - On deploy, walk `docs/` directory - Parse each markdown file: - Extract frontmatter for metadata (entity type, tags, relations) - Infer entity relationships from path conventions: `docs/containers/105-apps.md` → relationship to `entity:lxc:apps` - Extract cross-references (markdown links) → relationships - Create knowledge entities in the `entities` table (type = `document`, `runbook`, `investigation`, etc.) with `documents` / `procedure-for` edges to infrastructure entities - Idempotent — safe to re-run on every deploy **Agent access:** - MCP tool `search_knowledge(query)` — full-text search on knowledge entities - MCP tool `get_entity_knowledge(entity_id)` — all docs related to an entity - MCP tool `get_relations(entity_id)` — graph traversal (blast radius, dependencies) ### 7. Hermes agent container (`compose/hermes/`) - Hermes Agent runtime in a Docker container, gateway mode - Config: `hermes/config.yaml` (providers, models, gateway port) - Persona: `hermes/SOUL.md` (homelab-specific) - Skills: `hermes/skills/homelab-ops/SKILL.md` — how to use the Oikos API, classify actions, request approvals, query the knowledge graph - Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid) - Gateway port 8092 — workstations connect remotely - Hermes data volume for persistent state ### 8. Infisical secrets migration Same as rev 1: - Stand up Infisical in the Docker stack - Migrate SOPS secrets (one-time decrypt + import) - Wire all Go services to Infisical via machine identity - Retire SOPS + age keys ### 9. Notifier — Go (`cmd/notifier/`, `internal/notifier/`) **Interface (Go):** ```go type Notifier interface { SendAlert(ctx context.Context, signal Signal) error SendApprovalRequest(ctx context.Context, approval Approval) error ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error) } ``` **Matrix implementation:** - Sends alerts to `@dtoro:avispero` via Synapse (LXC 118) - Approval requests as messages with ✅/❌ reactions - Listens for reactions to record decisions - Pluggable — future implementations (webhook, email) register via config ### 10. Docker build + deploy pipeline - Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch) - `docker compose build` from repo root - Gitea webhook on push to `main` → deploy script on mac-mini - Deploy script: `git pull && go build ./... && docker compose build && docker compose up -d` - Seed ingest runs as part of the API startup (idempotent) - Health checks on each service ### 11. Ingress re-point (Caddy) - `dtoro/caddy-conf`: point `mcp.hubris.network` + `oikos.hubris.network` → mac-mini mesh IP :8090 (API) - Future: `hermes.hubris.network` → mac-mini:8092 - DNS and public URLs unchanged ### 12. mac-mini host setup - New directory: `~/oikos-os/` — repo clone + `docker compose` working dir - Existing `/opt/homelab-context/` stays untouched until cleanup phase - Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys - Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone ### 13. Decommission apps/105 (deferred) Keep apps/105 running as fallback. Cutover checklist when ready: 1. Verify Docker OS serves all traffic 2. `systemctl disable --now` Oikos services on apps/105 3. Remove old checkouts 4. Update Caddy backends exclusively to mac-mini 5. Remove/retarget Gitea webhooks ## Phasing **Phase 0 — Ontology design (no code):** - Finalize entity types, relationship types, lifecycles - Write seed manifests (`seeds/ontology.yaml`, `seeds/inventory.yaml`, `seeds/policy.yaml`) - Review diagrams with operator **Phase 1 — Foundation (Go + DB):** - Go module setup, project structure - PostgreSQL migrations (all 5) - Seed ingest pipeline (YAML → DB) - sqlc queries for core operations - Import existing signals/ledger data **Phase 2 — API (Go):** - Gin server with REST routes - MCP protocol adapter - Policy enforcement middleware - Knowledge graph ingestion from docs/ **Phase 3 — Control loop (Go):** - Scheduler (Observe) — probes, signals, state snapshots - Actuator (Act) — classify, execute, verify - Learning engine — feedback, patterns, skills **Phase 4 — Agent (Hermes container):** - Hermes Docker image, gateway config - Homelab skills - Connect from workstation, verify MCP + SSH **Phase 5 — Secrets (Infisical):** - Stand up Infisical, migrate SOPS, wire services **Phase 6 — Deploy + cutover:** - Docker Compose, Gitea webhook, Caddy re-point - End-to-end verification - Stop apps/105, clean up mac-mini ## Reuse (logic carried over, rewritten in Go) | Existing Python | What it becomes in Go | |---|---| | `oikos/decide.py` | `internal/policy/classify.go` — same scoring logic, reads from DB + patterns | | `oikos/signal.py` | `internal/scheduler/signal.go` — same lifecycle, DB-backed | | `oikos/approve.py` | `internal/policy/approve.go` — same grant lifecycle, DB-backed | | `oikos/ledger.py` | `internal/db/` — ledger_entries table + sqlc queries | | `oikos/policy.py` + `policy.yaml` | `internal/policy/` + `seeds/policy.yaml` → DB tables | | `oikos/drift.py` | `internal/scheduler/` — drift detection, writes signals to DB | | `oikos/relations.py` | `internal/ontology/graph.go` — SQL graph traversal | | `oikos/report.py` | `internal/api/routes/` — report endpoints, reads from DB | | `mcp/server.py` | `internal/api/mcp.go` — MCP adapter on top of DB | | `bin/homelab` logic | `internal/api/routes/` — same operations, REST interface | | `oikos/scheduler.py` | `internal/scheduler/` — same probes, goroutines for concurrency | | `oikos/approve.py` Matrix delivery | `internal/notifier/matrix.go` | | *(new)* | `internal/learning/` — feedback, patterns, skills (no Python equivalent) | ## Risks / trade-offs - **Go rewrite** — the existing Python code (~4400 lines) is replaced. The logic and design patterns carry over, but it's a full rewrite. Mitigated by the fact that the Python code is well-documented and the Go structure mirrors it. - **Hermes in Docker** — agent's world is the container. SSH access is the bridge. Hybrid approach (mounted keys now, actuator gateway later). - **PostgreSQL as SPOF** — mitigated by Docker volume persistence + automated `pg_dump` backups (the scheduler can do this once running). - **Learning model cold start** — no patterns/skills exist initially. The agent starts cautious (escalates everything), accumulates feedback, and gradually becomes more autonomous as patterns validate. This is by design — trust is earned. - **Infisical bootstrapping** — SOPS coexists during transition. Keep SOPS as fallback. - **Multi-agent concurrency** — `SELECT ... FOR UPDATE SKIP LOCKED` prevents two actuator passes from acting on the same signal. - **Ontology evolution** — as the homelab changes, entity types and relationship types need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy. ## Verification (end to end) 1. **Ontology:** `seeds/ontology.yaml` ingested — `SELECT * FROM entity_types` shows all types across 3 layers; lifecycle definitions match the state machine diagrams. 2. **DB:** `docker compose up postgres` — all 5 migrations applied; seed ingest populates entities + relationships from `inventory.yaml`. 3. **API:** `curl http://localhost:8090/api/v1/entities?type=service` returns the fleet; MCP `list_services` works via the same endpoint. 4. **Scheduler:** trigger a probe pass — signals in DB, state snapshots written. 5. **Actuator:** raise a test `service-down` signal → actuator classifies → auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated. 6. **Learning:** after N executions of the same (entity_type, action), a pattern appears with confidence score; after enough evidence, a skill is created. 7. **Classifier with learning:** set `autonomy.auto_act: reversible_low` → next similar signal: classifier checks pattern confidence → auto-acts if high, escalates if low. Kill-switch (`auto_act: off`) → always escalates. 8. **Hermes:** connect from another workstation → agent responds, queries API via MCP, can SSH to hubris. 9. **Secrets:** Infisical running, Go services fetch secrets, SOPS files removed. 10. **Deploy:** `git push` → Gitea webhook → `docker compose build + up -d` → changes live, seed ingest syncs any YAML changes to DB. 11. **Knowledge:** MCP `search_knowledge("caddy")` returns docs linked to `entity:service:caddy`. 12. **Cutover:** stop apps/105, verify production traffic only from Docker OS. ## Out of scope (for now) - Oikos Console web UI (deferred — built later on top of the API) - Multi-node deployment (designed for, not implemented) - Vector embeddings / semantic search (structured graph only for now) - SSH-key-signed approval requests - Prometheus / trend signals - Actuator gateway pattern (Phase 2 of hybrid — start with mounted SSH) - Automated skill extraction via LLM (patterns extracted statistically for now; LLM-assisted skill refinement is a future enhancement)