# Plan: Oikos — Docker-based agentic homelab OS on mac-mini **Status:** Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first design with systems modeling, DB-native config (inventory/ontology/policy as graph metadata, not YAML files), and a feedback loop where the agent learns from execution. ## Vision Convert this repo into a **Docker-based agentic homelab OS** written in **Go**. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors: - **Operator** (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix. - **Agent** — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and **learns from every action** to improve over time. All OS services run in Docker containers on mac-mini. The OS is deployed by a git push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path to multi-node later. ## Decisions (from operator Q&A, 2026-07-06) | Question | Decision | |---|---| | Repo structure | One repo, reorganize internally. | | Language | **Go** — compiled, type-safe, small containers, goroutines for concurrent probes. | | Hermes runtime | Runs inside Docker as part of the OS stack (gateway mode). | | Agent → homelab access | Hybrid — mounted SSH keys now, actuator gateway built incrementally. | | Data storage | PostgreSQL. Inventory/ontology/policy become DB-native graph metadata. | | Config (inventory/ontology/policy) | **DB-native** — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend. | | Knowledge/context | Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data. | | Operator interface | Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks. | | Host | mac-mini for now, designed to scale later. | | Deploy | Git push → Gitea webhook → Docker rebuild + restart. | | `bin/homelab` CLI | Replaced by an API. CLI becomes a thin Go client that calls the OS API. | | Secrets | Migrate from SOPS+age to Infisical. | | Matrix | Keep for now, abstract the notification layer for future channels. | | MCP server | Merged into the unified API — one Go service, REST + MCP interfaces. | | Agent type | Hermes core + custom homelab skills. | | Feedback loop | **Agent learns from execution** — outcomes feed back as patterns and skills that inform future decisions. | | Ontology | **Developed first**, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles. | | apps/105 | Keep running as fallback until Docker OS is proven. | | mac-mini cleanup | Start fresh in a new directory, clean up old artifacts later. | ## Ontology — the systems model The ontology is the foundational layer of the OS. It defines what exists, how things connect, how they change over time, and how the OS learns about them. It is stored **in the database** as metadata (entity types, relationship types, lifecycle definitions). YAML seed files bootstrap it on first deploy; after that, the DB is authoritative and editable via API. ### Design principles (systems modeling) 1. **Three layers** — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's own behavior + learning). Dependencies flow upward: Cognition depends on Governance depends on Infrastructure. 2. **Everything is an entity** — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, changes, skills) are first-class entities, not second-class records. 3. **Typed relationships with cardinality** — edges carry semantics. `hosts` is one-to-many; `depends-on` is many-to-many; `documents` is one-to-one. The graph is queryable for blast radius, dependency chains, and knowledge lookup. 4. **Lifecycles are state machines** — every entity type has a lifecycle. Infrastructure entities move through `planned → active → destroyed`. Operational entities have their own lifecycles (signals, approvals, patterns, skills). Transitions can require preconditions. 5. **Policy is attached to the ontology** — risk classes and approval rules link to entity types and actions. The policy IS part of the model, not a separate file. 6. **Learning is modeled** — executions produce outcomes, outcomes accumulate into patterns, patterns refine skills, skills inform future decisions. This is an explicit, queryable part of the graph. ### Layer map ```mermaid graph TB cognition -- "observes, acts on, learns about" --> infra governance -- "governs access to" --> infra governance -- "constrains" --> cognition cognition -- "creates + refines" --> governance subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"] direction LR OBS["Observation\nsignal, state-snapshot"] DEC["Decision\nclassification, risk-assessment"] ACT["Action\nexecution, verification"] GOV["Governance\napproval-request, approval-decision"] KNOW["Knowledge\ndocument, runbook"] LEARN["Learning\npattern, skill, feedback"] end subgraph governance["Layer 2 — Governance (who controls what)"] direction LR IDENT["Identity\nperson, agent, identity-provider"] SEC["Secrets\nsecret, key, access-grant"] POL["Policy\nrisk-class, approval-rule, autonomy-setting"] end subgraph infra["Layer 1 — Infrastructure (the managed world)"] direction LR PHYS["Physical\nsite, machine, ups, sensor"] COMP["Compute\nmachine, vm, container\n(lxc, docker)"] NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"] STOR["Storage\nstorage-pool, volume,\nmount, backup-target"] SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"] end ``` ### Block definition diagram (SysML BDD) Uses Mermaid class diagram syntax following SysML BDD conventions: - `«abstract»` = abstract block (cannot be instantiated) - `<\|--` = generalization (is-a) - `*--` = composition (whole-part, lifecycle dependency) - `o--` = aggregation (whole-part, independent lifecycle) - `--` = association (typed link) - Multiplicity: `"1"`, `"0..1"`, `"1..*"`, `"0..*"`, `"*"` **Infrastructure layer — compute, storage, network:** ```mermaid classDiagram class ComputeEntity { <> +state lifecycle +attributes jsonb } class Machine { +cpu_arch +ram_gb } class VirtualMachine { +vcpus +memory_mb +disk_gb } class Container { <> +runtime } class LXC { +pve_id +rootfs } class DockerContainer { +image +compose_stack } class ProxmoxHost { +pve_version +cluster_member } class StandaloneServer { +hypervisor } class Workstation { +os +user } class Appliance { +vendor +model } class Hypervisor { +type +version } ComputeEntity <|-- Machine ComputeEntity <|-- VirtualMachine ComputeEntity <|-- Container Machine <|-- ProxmoxHost Machine <|-- StandaloneServer Machine <|-- Workstation Machine <|-- Appliance Container <|-- LXC Container <|-- DockerContainer Machine "1" *-- "0..1" Hypervisor : runs Hypervisor "1" o-- "0..*" VirtualMachine : hosts Hypervisor "1" o-- "0..*" Container : hosts class StoragePool { +type lvm, zfs, nfs +capacity_gb } class Volume { +name +size_gb } class Mount { +mount_point +options } StoragePool "1" *-- "0..*" Volume : contains ComputeEntity "1" o-- "0..*" Mount : has Mount "0..*" --> "1" Volume : mounts class NetworkInterface { +mac +ip } class Network { <> } class LAN { +subnet} class Mesh { +provider} class VLAN { +tag} ComputeEntity "1" *-- "0..*" NetworkInterface : has NetworkInterface "0..*" --> "1" Network : connects-to Network <|-- LAN Network <|-- Mesh Network <|-- VLAN ``` **Software + services layer:** ```mermaid classDiagram class Service { +port +health_url +risk_notes } class Application { +version +config } class ConfigRepo { +url +branch } class DeployPipeline { +trigger +target_path } class IngressRoute { +pattern +upstream } class Certificate { +issuer +expires } class DNSZone { +zone } class DNSRecord { +name +record_type +value } ComputeEntity "1" o-- "0..*" Service : provides Service "1" *-- "0..*" Application : runs Service "0..1" --> "0..1" ConfigRepo : configured-by DeployPipeline "0..*" --> "1" Service : deploys-to IngressRoute "0..*" --> "1" Service : routes-to IngressRoute "0..*" --> "0..1" Certificate : secured-by IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by Service "0..*" --> "0..*" Service : depends-on DNSZone "1" *-- "0..*" DNSRecord : contains DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to ``` **Cognition layer — operations + learning:** ```mermaid classDiagram class Signal { +kind +severity +state lifecycle +evidence +recommended_action } class Execution { +status +result +duration_ms +verified } class Feedback { +outcome +observation +lesson } class Pattern { +confidence +evidence_count +status } class Skill { +procedure +version +success_rate } class Classification { +risk +route +reasoning } class Approval { +status +ttl } class Document { +title +content +source_path } class Runbook { +steps +risk_class +verification } ComputeEntity "1" o-- "0..*" Signal : monitored-by Signal "0..1" --> "0..*" Execution : triggers Classification "1" --> "0..1" Execution : precedes Execution "1" *-- "0..1" Feedback : produces Feedback "0..*" --> "0..*" Pattern : contributes-to Pattern "0..*" --> "0..1" Skill : informs Skill "0..1" --> "0..*" Classification : guides Execution "0..1" --> "0..1" Approval : requires Agent "0..*" --> "0..*" Execution : performs Person "0..1" --> "0..*" Approval : decides Person "0..1" --> "0..*" Agent : owns Entity "1" o-- "0..*" Document : documented-by Entity "1" o-- "0..*" Runbook : procedure-for ``` ### Design notes — validated assumptions **Can VMs mount storage pools?** Yes. In Proxmox, VMs have virtual disks on storage pools (LVM, ZFS, NFS). LXCs have bind mounts and mount points. Both use the same storage pools. The model reflects this: `ComputeEntity` (abstract) has `Mount` edges to `Volume`, regardless of whether the compute entity is a VM, LXC, or machine. **Not every machine is a Proxmox host.** The current homelab has: 2 Proxmox hosts, 2 workstations, 1 external VPS, 2 VMs, 19 LXCs. The model uses `Machine` as the base type with specializations (`ProxmoxHost`, `StandaloneServer`, `Workstation`, `Appliance`). Only `ProxmoxHost` runs PVE; `StandaloneServer` could run KVM/libvirt; `Workstation` runs desktop OS + optionally Docker. A `Hypervisor` is software that runs on a `Machine` and hosts VMs/containers — it's not always present (workstations and appliances may not have one). **Docker containers are first-class compute entities.** The OS itself runs in Docker containers, and services like Jellyfin's MariaDB sidecar run in Docker within LXCs. `DockerContainer` is a specialization of `Container` with `image`, `compose_stack` attributes. This lets the OS model its own infrastructure. **Services run on any compute entity.** The homelab has services on LXCs (caddy, gitea), VMs (zimaos, haos), workstations (mac-mini will run the OS), and external hosts (authentik on netbird-vps). The `provides` relationship is from `ComputeEntity` (abstract), not from a specific compute type. **Documents and runbooks describe any entity.** The old ER diagram had `Document → Service` only. In practice, docs describe hosts, containers, network infrastructure, storage, and investigations. The model uses `Entity` (the root abstract type) for `documented-by` and `procedure-for`, so any entity can have docs and runbooks. ### Infrastructure lifecycle ```mermaid stateDiagram-v2 [*] --> planned : operator creates entity planned --> provisioning : IP reserved, storage chosen, doc stub provisioning --> active : mesh joined, health-check answering, doc complete active --> migrating : preflight + backup verified migrating --> active : post-verify, caddy checked, mounts checked active --> deprecated : replacement live or role retired deprecated --> destroyed : backups verified, secrets revoked, ingress removed destroyed --> [*] : archaeology entry recorded note right of deprecated Complete only when zero inbound edges remain (no depends-on, no routes-to) end note ``` ### Signal lifecycle ```mermaid stateDiagram-v2 [*] --> raised : scheduler probe or agent finding raised --> acknowledged : agent or operator sees it raised --> muted : operator suppresses (TTL) acknowledged --> acting : actuator starts execution acting --> resolved : action succeeded, verification passed acting --> raised : action failed, re-escalated raised --> resolved : condition cleared (auto-resolve) muted --> raised : TTL expired resolved --> [*] ``` ### Execution + learning lifecycle ```mermaid stateDiagram-v2 [*] --> proposed : signal + recommended_action proposed --> approved : operator approves (if gated) proposed --> auto_approved : risk class allows auto-act approved --> executing : actuator runs auto_approved --> executing : actuator runs executing --> verified : verification command passed executing --> failed : execution or verification failed executing --> timed_out : exceeded duration limit failed --> rolled_back : rollback procedure executed verified --> [*] : feedback recorded, pattern updated failed --> [*] : feedback recorded, pattern updated rolled_back --> [*] : feedback recorded, pattern updated timed_out --> [*] : feedback recorded, pattern updated ``` ### Feedback / learning model — the cognition loop ```mermaid flowchart TB subgraph observe["Observe"] SIG["Signal raised\n(service down, disk full, drift)"] end subgraph decide["Decide"] CLASS["Classification\nrisk × blast radius × confidence"] SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"] CLASS --> SKILL_LOOKUP end subgraph act["Act"] EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"] VERIFY["Verification\ncheck if action succeeded"] EXEC --> VERIFY end subgraph learn["Learn (feedback loop)"] OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"] FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"] PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"] SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"] OUTCOME --> FEEDBACK FEEDBACK --> PATTERN PATTERN --> SKILL_REFINE end SIG --> CLASS SKILL_LOOKUP --> EXEC VERIFY --> OUTCOME SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP subgraph escalation["Escalation path"] APPROVAL["Approval request\n→ Matrix ✅/❌"] end CLASS -- "needs approval" --> APPROVAL APPROVAL -- "approved" --> EXEC APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"] ``` ### Policy model — how risk attaches to entities ```mermaid flowchart LR subgraph ontology["Ontology (DB metadata)"] ET["entity_types\nhost, service, lxc, vm, ..."] RT["relationship_types\nhosts, provides, depends-on, ..."] LD["lifecycle_defs\nstates + transitions"] end subgraph policy["Policy (DB records)"] RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"] RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"] AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"] end subgraph instances["Instances (DB data)"] ENT["entities\nhost:hubris, service:caddy, ..."] REL["relationships\nhubris hosts lxc:apps"] end ET --> RULES RC --> RULES RULES --> ENT AUTO --> ENT ET --> ENT RT --> REL LD --> ENT ``` ## Target architecture ### Container stack on mac-mini ```mermaid graph TB subgraph mac-mini["mac-mini — Docker host, always-on"] subgraph services["Docker Compose — Go binaries"] PG["PostgreSQL 16\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata"] INF["Infisical\n(secrets manager)"] HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"] API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events\n\nPolicy enforcement + risk\nclassification + ledger"] SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• SSH to hubris/strong"] NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"] end DEPLOY["Gitea webhook →\ndocker compose build + up -d"] end HERMES -- MCP --> API API --> PG SCHED --> PG SCHED -- SSH --> HUBRIS SCHED -- SSH --> STRONG API -- "escalate" --> NOTIFIER NOTIFIER -- "alerts + approvals" --> MATRIX DEPLOY -- "rebuild" --> services subgraph external["External"] CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"] APPS["apps/105 (fallback)"] HUBRIS["hubris (PVE)"] STRONG["strong (PVE)"] GITEA["Gitea (LXC 104)"] MATRIX["Matrix (LXC 118)"] WS["Any workstation\nHermes remote → gateway"] end CADDY -- reverse_proxy --> API GITEA -- webhook --> DEPLOY WS -- "Hermes gateway" --> HERMES ``` ### OODA loop — with learning feedback ```mermaid flowchart LR OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"] ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"] DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"] DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"] APPROVE -- "approved" --> ACT ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"] LEARN -. "improves confidence" .-> DECIDE LEARN --> OBSERVE ``` ### Deploy flow ```mermaid flowchart LR DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"] GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"] MACMINI -- "git pull" --> REPO["repo clone"] MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"] STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"] ``` ## DB-native configuration The three YAML files — `inventory.yaml`, `ontology.yaml`, `policy.yaml` — become **seed manifests**. They bootstrap the DB on first deploy. After that, the DB is the runtime source of truth, editable via the API. A future frontend can edit all three directly. ### How it works ```mermaid flowchart TB subgraph seeds["Seed manifests (git-tracked, YAML)"] ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"] INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"] POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"] end subgraph db["PostgreSQL (runtime source of truth)"] META["entity_types table\nrelationship_types table\nlifecycle_defs table"] INST["entities table\nrelationships table"] POLDB["policies table\nrisk_classes table\nautonomy_settings table"] end INGEST["Seed ingest (on deploy)\nidempotent upsert"] ONTO_YAML --> INGEST --> META INV_YAML --> INGEST --> INST POL_YAML --> INGEST --> POLDB API_EDIT["API edits\n(POST/PUT/PATCH)"] API_EDIT --> META API_EDIT --> INST API_EDIT --> POLDB EXPORT["Export to YAML\n(for DR / version control)"] META --> EXPORT INST --> EXPORT POLDB --> EXPORT ``` ### Why DB-native - **Querying** — the agent can ask "what services depend on authentik?" as a graph query, not a YAML parse. Blast-radius walks are SQL, not file reads. - **Mutation** — adding a service, updating a lifecycle state, changing a policy rule are DB transactions with audit trail, not file edits + git commits. - **Consistency** — the ontology, inventory, and policy are always in sync (same DB, same transaction). No drift between what the YAML says and what the runtime sees. - **Future frontend** — a UI can edit entities, relationships, and policies directly via the API. No need to generate/edit YAML files. - **Version control** — seed YAML files are still git-tracked for bootstrap and DR. The API can export the current DB state back to YAML for commit. ## Repo layout (Go project) ``` / # repo root ├── docker-compose.yml # the OS stack definition ├── Makefile # build, test, deploy targets ├── go.mod # Go module definition ├── go.sum ├── cmd/ # binary entrypoints (one per service) │ ├── api/ # Oikos API server │ │ └── main.go │ ├── scheduler/ # Observe + Act loop │ │ └── main.go │ └── notifier/ # Notification service │ └── main.go ├── internal/ # private packages (not importable) │ ├── db/ # database layer │ │ ├── queries/ # sqlc SQL queries │ │ ├── models.go # generated Go types │ │ └── db.go # connection pool, migrations │ ├── ontology/ # ontology types + meta-schema │ │ ├── types.go # EntityType, RelationshipType, LifecycleDef │ │ ├── graph.go # graph traversal (blast radius, dependencies) │ │ └── ingest.go # YAML seed → DB ingest │ ├── api/ # HTTP + MCP server │ │ ├── server.go # Gin app setup │ │ ├── routes/ # REST handlers │ │ │ ├── hosts.go │ │ │ ├── services.go │ │ │ ├── signals.go │ │ │ ├── approvals.go │ │ │ ├── exec.go │ │ │ └── knowledge.go │ │ └── mcp.go # MCP protocol adapter (JSON-RPC over SSE) │ ├── policy/ # risk classification + approval │ │ ├── classify.go # risk × blast × confidence │ │ ├── approve.go # approval request + grant lifecycle │ │ └── autonomy.go # kill-switch, never-auto-act list │ ├── scheduler/ # Observe stage │ │ ├── probe.go # HTTP health, disk, drift │ │ └── signal.go # raise/resolve signals in DB │ ├── actuator/ # Act stage │ │ ├── act.go # read signals, classify, execute or escalate │ │ ├── execute.go # SSH execution + verification │ │ └── guard.go # loop-guard, retry caps │ ├── learning/ # feedback loop (the learning model) │ │ ├── feedback.go # record outcome + lesson from execution │ │ ├── pattern.go # extract/validate patterns from feedback │ │ └── skill.go # create/refine skills from patterns │ ├── notifier/ # notification abstraction │ │ ├── notifier.go # interface │ │ └── matrix.go # Matrix implementation │ └── config/ # config loading (env, files) │ └── config.go ├── migrations/ # SQL migrations (golang-migrate format) │ ├── 001_ontology.up.sql # meta-schema (entity_types, relationship_types, ...) │ ├── 001_ontology.down.sql │ ├── 002_instances.up.sql # entities, relationships │ ├── 003_operations.up.sql # signals, approvals, executions, feedback │ ├── 004_learning.up.sql # patterns, skills │ └── 005_policy.up.sql # policies, risk_classes, autonomy ├── seeds/ # YAML seed manifests (bootstrap + DR) │ ├── ontology.yaml # entity types, relationship types, lifecycles │ ├── inventory.yaml # entity instances (hosts, services, etc.) │ └── policy.yaml # risk classes, approval rules, autonomy ├── docs/ # narrative docs (ingested into knowledge graph) │ ├── containers/ │ ├── hosts/ │ ├── infrastructure/ │ └── investigations/ ├── hermes/ # Hermes agent config + skills │ ├── config.yaml │ ├── SOUL.md │ └── skills/ │ └── homelab-ops/ │ └── SKILL.md ├── compose/ # Docker build contexts │ ├── api/Dockerfile │ ├── scheduler/Dockerfile │ ├── hermes/Dockerfile │ └── postgres/init.sql └── scripts/ # utility scripts ├── migrate-sops.sh # one-time SOPS → Infisical migration └── import-legacy.sh # import existing signals/ledger JSONL ``` ## Database schema The schema is the ontology made concrete. Five migration groups, each adding a layer. ### Migration 1: Ontology meta-schema ```sql -- The meta-graph: defines what entity types and relationship types can exist. -- This IS the ontology, stored in the DB, editable via API. CREATE TABLE entity_types ( name TEXT PRIMARY KEY, -- 'host', 'service', 'signal', 'pattern' domain TEXT NOT NULL, -- 'physical', 'compute', 'network', ... layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition' description TEXT, lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle) attribute_schema JSONB, -- JSON Schema for validating entity.attributes status TEXT NOT NULL DEFAULT 'active', -- 'active', 'deprecated' (no hard delete while instances exist) created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE TABLE relationship_types ( name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on' inverse TEXT, -- 'runs-on', 'provided-by' source_type TEXT REFERENCES entity_types(name), target_type TEXT REFERENCES entity_types(name), cardinality TEXT NOT NULL, -- 'one-to-one', 'one-to-many', 'many-to-many' description TEXT, created_at TIMESTAMPTZ DEFAULT now() ); CREATE TABLE lifecycle_defs ( id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', 'execution', ... states TEXT[] NOT NULL, -- ordered states default_state TEXT NOT NULL, transitions JSONB NOT NULL, -- {"from": {"to": {"requires": [...]}}} created_at TIMESTAMPTZ DEFAULT now() ); ``` ### Migration 2: Entity instances (the inventory graph) ```sql -- Entity instances — the actual hosts, services, signals, patterns, etc. -- This replaces inventory.yaml as the runtime source of truth. CREATE TABLE entities ( id TEXT PRIMARY KEY, -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001' type TEXT NOT NULL REFERENCES entity_types(name), name TEXT NOT NULL, -- 'hubris', 'caddy', 'disk-threshold' state TEXT, -- lifecycle state (e.g. 'active', 'raised') attributes JSONB NOT NULL DEFAULT '{}', -- type-specific data (IP, mesh addr, port, ...) parent_id TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host) created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE INDEX idx_entities_type ON entities(type); CREATE INDEX idx_entities_state ON entities(state); CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes); -- Relationship instances — the typed edges of the graph CREATE TABLE relationships ( source_id TEXT NOT NULL REFERENCES entities(id), target_id TEXT NOT NULL REFERENCES entities(id), type TEXT NOT NULL REFERENCES relationship_types(name), attributes JSONB, created_at TIMESTAMPTZ DEFAULT now(), PRIMARY KEY (source_id, target_id, type) ); CREATE INDEX idx_rel_source ON relationships(source_id); CREATE INDEX idx_rel_target ON relationships(target_id); CREATE INDEX idx_rel_type ON relationships(type); -- Recursive graph traversal function (blast radius, dependency chains) CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3) RETURNS TABLE(entity_id TEXT, depth INT) AS $$ WITH RECURSIVE walk AS ( SELECT start_id::TEXT AS entity_id, 0::INT AS depth UNION SELECT r.target_id::TEXT, w.depth + 1 FROM relationships r JOIN walk w ON r.source_id = w.entity_id WHERE w.depth < max_depth ) SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id; $$ LANGUAGE sql STABLE; ``` ### Migration 3: Operations (signals, approvals, ledger, state) ```sql -- Signals — now entities in the graph, with a dedicated table for indexed querying -- (the entity row is the canonical record; this table is a fast lookup) CREATE TABLE signals ( entity_id TEXT PRIMARY KEY REFERENCES entities(id), kind TEXT NOT NULL, severity TEXT NOT NULL, -- info, warning, critical target_entity_id TEXT REFERENCES entities(id), -- the infrastructure entity this is about evidence TEXT, likely_cause TEXT, recommended_action JSONB, verification TEXT, state TEXT NOT NULL DEFAULT 'raised', mute_until TIMESTAMPTZ, created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); CREATE INDEX idx_signals_state ON signals(state); CREATE INDEX idx_signals_target ON signals(target_entity_id); CREATE INDEX idx_signals_severity ON signals(severity); -- Approvals CREATE TABLE approvals ( id TEXT PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), -- entity to act on action TEXT NOT NULL, risk_class TEXT NOT NULL, status TEXT NOT NULL DEFAULT 'pending', -- pending, approved, denied, expired ttl INTERVAL NOT NULL DEFAULT '1 hour', decided_at TIMESTAMPTZ, decided_by TEXT REFERENCES entities(id), -- person entity confirmation_phrase TEXT ); -- Change ledger (high-level record, links to execution for detail) CREATE TABLE ledger_entries ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), action TEXT NOT NULL, risk_class TEXT NOT NULL, result TEXT, -- ok, failed, escalated approval_id TEXT REFERENCES approvals(id), execution_id INTEGER, -- FK to executions (migration 4) agent_id TEXT REFERENCES entities(id), notes TEXT ); -- State snapshots (replaces oikos/state.json) CREATE TABLE state_snapshots ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), entity_id TEXT REFERENCES entities(id), health TEXT, -- healthy, degraded, down, unknown data JSONB ); ``` ### Migration 4: Learning model (executions, feedback, patterns, skills) ```sql -- Executions — detailed record of each action the OS performs CREATE TABLE executions ( id SERIAL PRIMARY KEY, ts TIMESTAMPTZ DEFAULT now(), signal_entity_id TEXT REFERENCES entities(id), -- signal that triggered this target_entity_id TEXT REFERENCES entities(id), -- entity acted upon action TEXT NOT NULL, risk_class TEXT NOT NULL, approval_id TEXT REFERENCES approvals(id), agent_id TEXT REFERENCES entities(id), -- who/what executed skill_id TEXT REFERENCES entities(id), -- skill used (if any) status TEXT NOT NULL DEFAULT 'queued', -- queued, running, completed, failed, timed-out result JSONB, -- detailed result data duration_ms INTEGER, verified BOOLEAN DEFAULT false, started_at TIMESTAMPTZ, completed_at TIMESTAMPTZ ); CREATE INDEX idx_exec_target ON executions(target_entity_id); CREATE INDEX idx_exec_status ON executions(status); CREATE INDEX idx_exec_action ON executions(action); -- Feedback — what was learned from an execution CREATE TABLE feedback ( id SERIAL PRIMARY KEY, execution_id INTEGER REFERENCES executions(id), ts TIMESTAMPTZ DEFAULT now(), outcome TEXT NOT NULL, -- success, failure, partial, unexpected observation TEXT, -- what happened vs what was expected lesson TEXT, -- extractable lesson unexpected_side_effects TEXT[], tags TEXT[] ); CREATE INDEX idx_feedback_execution ON feedback(execution_id); CREATE INDEX idx_feedback_outcome ON feedback(outcome); -- Patterns — generalized rules extracted from accumulated feedback CREATE TABLE patterns ( id TEXT PRIMARY KEY, -- 'pat-2026-07-06-001' ts TIMESTAMPTZ DEFAULT now(), entity_type TEXT REFERENCES entity_types(name), -- applies to this type action TEXT NOT NULL, -- 'restart', 'deploy', etc. pattern TEXT NOT NULL, -- 'service X recovers within 30s after restart' confidence REAL DEFAULT 0.5, -- 0.0 to 1.0 evidence_count INTEGER DEFAULT 1, -- how many executions support this success_count INTEGER DEFAULT 0, failure_count INTEGER DEFAULT 0, status TEXT DEFAULT 'hypothesized', -- hypothesized, validated, active, deprecated last_validated_at TIMESTAMPTZ ); CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action); CREATE INDEX idx_patterns_status ON patterns(status); -- Skills — codified procedures refined through feedback CREATE TABLE skills ( id TEXT PRIMARY KEY, -- 'skill-restart-service', 'skill-deploy-lxc' name TEXT NOT NULL, ts TIMESTAMPTZ DEFAULT now(), procedure TEXT NOT NULL, -- the codified steps (markdown or structured) applies_to TEXT REFERENCES entity_types(name), pattern_ids TEXT[], -- patterns that inform this skill status TEXT DEFAULT 'drafted', -- drafted, tested, active, refined, deprecated version INTEGER DEFAULT 1, success_rate REAL, -- rolling success rate last_used_at TIMESTAMPTZ ); CREATE INDEX idx_skills_type ON skills(applies_to); CREATE INDEX idx_skills_status ON skills(status); ``` ### Migration 5: Policy (DB-native risk + approval rules) ```sql -- Risk classes — the four-level safety model CREATE TABLE risk_classes ( name TEXT PRIMARY KEY, -- 'read_only', 'reversible_low', etc. description TEXT, approval_required TEXT NOT NULL DEFAULT 'none', -- none, operator, operator_confirmed ledger BOOLEAN DEFAULT false, autonomy_allowed BOOLEAN DEFAULT false -- can agent auto-act at this risk level? ); -- Approval rules — entity_type + action → risk_class + requirements CREATE TABLE approval_rules ( id SERIAL PRIMARY KEY, entity_type TEXT REFERENCES entity_types(name), -- applies to this entity type action TEXT NOT NULL, -- 'restart', 'deploy', 'destroy' risk_class TEXT NOT NULL REFERENCES risk_classes(name), autonomy_level TEXT NOT NULL DEFAULT 'auto', -- 'auto', 'escalate', 'never' scope_entity TEXT REFERENCES entities(id), -- optional: specific entity only created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now(), UNIQUE(entity_type, action) ); -- Autonomy settings — global kill-switch + per-entity overrides CREATE TABLE autonomy_settings ( key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.caddy' value TEXT NOT NULL, -- 'off', 'reversible_low', 'true', 'false' updated_at TIMESTAMPTZ DEFAULT now() ); ``` ## Workstreams ### 1. Ontology definition + seed manifests (`seeds/`, `internal/ontology/`) **Before any code**, finalize the ontology. The current `ontology.yaml` has 8 domains and 14 relationship types. The new ontology adds: - **Layer 3 entities** (cognition): signal, execution, feedback, pattern, skill, approval, classification, document, runbook - **New relationships**: `triggers`, `produces`, `contributes-to`, `informs`, `guides`, `precedes`, `performs`, `procedure-for`, `learned-from` - **Lifecycle definitions** for operational entities (signals, executions, patterns, skills) — not just infrastructure **Deliverables:** - `seeds/ontology.yaml` — entity types, relationship types, lifecycle definitions (seeded into `entity_types`, `relationship_types`, `lifecycle_defs` on deploy) - `seeds/inventory.yaml` — adapted from current `inventory.yaml` (seeded into `entities` + `relationships`) - `seeds/policy.yaml` — adapted from current `policy.yaml` (seeded into `risk_classes`, `approval_rules`, `autonomy_settings`) - `internal/ontology/ingest.go` — idempotent seed → DB ingest ### 2. Database layer + migrations (`migrations/`, `internal/db/`) - Write the 5 migrations above - Set up sqlc for type-safe Go database access - Connection pool, migration runner - Graph traversal queries (blast radius, dependency chains, knowledge lookup) - Import script for existing `signals/*.jsonl`, `ledger/*.jsonl` → DB ### 3. Unified API server — Go (`cmd/api/`, `internal/api/`) One Go binary (Gin web framework) exposing REST + MCP from the same codebase. **MCP interface** (`internal/api/mcp.go`): - JSON-RPC over SSE, compatible with Hermes MCP client - Tools: `get_host`, `list_services`, `search_knowledge`, `get_entity`, `get_relations`, `get_blast_radius`, `get_signal_history`, `get_ledger`, `get_patterns`, `get_skills`, `get_state_snapshot` - All read from PostgreSQL **REST interface** (`internal/api/routes/`): - `GET /api/v1/entities` — list entities (filter by type, state, domain) - `GET /api/v1/entities/{id}` — entity detail + relationships - `POST /api/v1/entities` — create entity (creates inventory entry) - `PATCH /api/v1/entities/{id}` — update entity (state transition, attributes) - `GET /api/v1/signals` — list signals (filter by state, severity, entity) - `POST /api/v1/signals/{id}/ack` — acknowledge - `POST /api/v1/signals/{id}/resolve` — resolve - `GET /api/v1/approvals` — pending approvals - `POST /api/v1/approvals/{id}/decide` — approve/deny (Authentik-gated) - `POST /api/v1/exec` — gated execution (classify → check approval → execute → feedback) - `GET /api/v1/patterns` — list patterns (filter by entity_type, action, status) - `GET /api/v1/skills` — list skills (filter by applies_to, status) - `GET /api/v1/knowledge/{entity_id}` — knowledge graph query - `GET /api/v1/knowledge/search?q=...` — search knowledge graph - `GET /api/v1/ontology` — list entity types, relationship types, lifecycles - `POST /api/v1/ontology/entity_types` — create entity type (extend the schema) - `WS /api/v1/events` — real-time stream (signals, approvals, executions, feedback) **Policy enforcement** (`internal/policy/classify.go`): - Every mutating endpoint classifies the action via the policy DB - Risk class → approval check → autonomy check - All mutations write to the ledger automatically **Auth:** - MCP interface: no auth (internal, container-to-container) - REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints - Internal: shared secret (Docker network) ### 4. Scheduler + Actuator — Go (`cmd/scheduler/`, `internal/scheduler/`, `internal/actuator/`) **Scheduler (Observe)** — Go service with goroutines for concurrent probes: - HTTP health probes (concurrent, with timeouts) - Disk usage probes (SSH to hubris/strong) - Drift detection (inventory vs live state) - Writes signals + state snapshots to DB - Runs on a 10-min ticker **Actuator (Act)** — Go service, the control loop: - Reads open signals with `recommended_action` - For each: classify via `internal/policy/classify.go` - **auto-act**: look up skill for (entity_type, action) → follow procedure → execute via SSH → verify → record execution → generate feedback → update patterns - **escalate**: create approval request → notify via Matrix → acknowledge signal - Loop-guard: check execution history per (entity, action) to cap auto-retries (`SELECT ... FOR UPDATE SKIP LOCKED` for concurrency safety) - Autonomy kill-switch: check `autonomy_settings` table ### 5. Learning engine — Go (`internal/learning/`) The feedback loop that makes the agent improve over time. **Feedback recording** (`internal/learning/feedback.go`): - After every execution, evaluate the outcome: - Did the verification command pass? → success - Did it fail? → failure - Did it partially work? → partial - Did something unexpected happen? → unexpected - Record a feedback entry with: outcome, observation (what happened vs expected), lesson (extractable insight), unexpected_side_effects **Pattern extraction** (`internal/learning/pattern.go`): - Periodically scan accumulated feedback for (entity_type, action) pairs - When N+ executions share a similar outcome, extract a pattern: - "Restarting service:X typically takes 15s and succeeds" - "Deploying to LXC:Y via webhook has 30% failure rate, retry helps" - Patterns start as `hypothesized`, move to `validated` after enough evidence, then `active` (used by the decision classifier) - Confidence score = success_count / evidence_count, adjusted by recency **Skill management** (`internal/learning/skill.go`): - When a pattern reaches `active` status with confidence > 0.7, create or refine a skill for that (entity_type, action) pair - Skills codify the best-known procedure (what steps to take, what to verify, expected duration, known failure modes) - Skills are versioned — each refinement increments the version - The actuator looks up skills before executing: if a skill exists, follow it; if not, use the default procedure and generate feedback for future pattern extraction **How the classifier uses learning** (`internal/policy/classify.go`): - Confidence scoring now checks patterns + skills, not just raw ledger history: - If a pattern exists for (entity_type, action) with high confidence → boost auto-act confidence - If patterns show frequent failures → lower confidence, escalate - If a skill exists → higher confidence (proven procedure available) - This is the closed loop: **execution → feedback → pattern → skill → classification → execution** (better informed each time) ### 6. Knowledge graph ingestion (`internal/ontology/ingest.go`) After the ontology is defined and the DB schema is in place: - On deploy, walk `docs/` directory - Parse each markdown file: - Extract frontmatter for metadata (entity type, tags, relations) - Infer entity relationships from path conventions: `docs/containers/105-apps.md` → relationship to `entity:lxc:apps` - Extract cross-references (markdown links) → relationships - Create knowledge entities in the `entities` table (type = `document`, `runbook`, `investigation`, etc.) with `documents` / `procedure-for` edges to infrastructure entities - Idempotent — safe to re-run on every deploy **Agent access:** - MCP tool `search_knowledge(query)` — full-text search on knowledge entities - MCP tool `get_entity_knowledge(entity_id)` — all docs related to an entity - MCP tool `get_relations(entity_id)` — graph traversal (blast radius, dependencies) ### 7. Hermes agent container (`compose/hermes/`) - Hermes Agent runtime in a Docker container, gateway mode - Config: `hermes/config.yaml` (providers, models, gateway port) - Persona: `hermes/SOUL.md` (homelab-specific) - Skills: `hermes/skills/homelab-ops/SKILL.md` — how to use the Oikos API, classify actions, request approvals, query the knowledge graph - Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid) - Gateway port 8092 — workstations connect remotely - Hermes data volume for persistent state ### 8. Infisical secrets migration Same as rev 1: - Stand up Infisical in the Docker stack - Migrate SOPS secrets (one-time decrypt + import) - Wire all Go services to Infisical via machine identity - Retire SOPS + age keys ### 9. Notifier — Go (`cmd/notifier/`, `internal/notifier/`) **Interface (Go):** ```go type Notifier interface { SendAlert(ctx context.Context, signal Signal) error SendApprovalRequest(ctx context.Context, approval Approval) error ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error) } ``` **Matrix implementation:** - Sends alerts to `@dtoro:avispero` via Synapse (LXC 118) - Approval requests as messages with ✅/❌ reactions - Listens for reactions to record decisions - Pluggable — future implementations (webhook, email) register via config ### 10. Docker build + deploy pipeline - Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch) - `docker compose build` from repo root - Gitea webhook on push to `main` → deploy script on mac-mini - Deploy script: `git pull && go build ./... && docker compose build && docker compose up -d` - Seed ingest runs as part of the API startup (idempotent) - Health checks on each service ### 11. Ingress re-point (Caddy) - `dtoro/caddy-conf`: point `mcp.hubris.network` + `oikos.hubris.network` → mac-mini mesh IP :8090 (API) - Future: `hermes.hubris.network` → mac-mini:8092 - DNS and public URLs unchanged ### 12. mac-mini host setup - New directory: `~/oikos-os/` — repo clone + `docker compose` working dir - Existing `/opt/homelab-context/` stays untouched until cleanup phase - Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys - Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone ### 13. Decommission apps/105 (deferred) Keep apps/105 running as fallback. Cutover checklist when ready: 1. Verify Docker OS serves all traffic 2. `systemctl disable --now` Oikos services on apps/105 3. Remove old checkouts 4. Update Caddy backends exclusively to mac-mini 5. Remove/retarget Gitea webhooks ## Phasing **Phase 0 — Ontology design (no code):** - Finalize entity types, relationship types, lifecycles - Write seed manifests (`seeds/ontology.yaml`, `seeds/inventory.yaml`, `seeds/policy.yaml`) - Review diagrams with operator **Phase 1 — Foundation (Go + DB):** - Go module setup, project structure - PostgreSQL migrations (all 5) - Seed ingest pipeline (YAML → DB) - sqlc queries for core operations - Import existing signals/ledger data **Phase 2 — API (Go):** - Gin server with REST routes - MCP protocol adapter - Policy enforcement middleware - Knowledge graph ingestion from docs/ **Phase 3 — Control loop (Go):** - Scheduler (Observe) — probes, signals, state snapshots - Actuator (Act) — classify, execute, verify - Learning engine — feedback, patterns, skills **Phase 4 — Agent (Hermes container):** - Hermes Docker image, gateway config - Homelab skills - Connect from workstation, verify MCP + SSH **Phase 5 — Secrets (Infisical):** - Stand up Infisical, migrate SOPS, wire services **Phase 6 — Deploy + cutover:** - Docker Compose, Gitea webhook, Caddy re-point - End-to-end verification - Stop apps/105, clean up mac-mini ## Reuse (logic carried over, rewritten in Go) | Existing Python | What it becomes in Go | |---|---| | `oikos/decide.py` | `internal/policy/classify.go` — same scoring logic, reads from DB + patterns | | `oikos/signal.py` | `internal/scheduler/signal.go` — same lifecycle, DB-backed | | `oikos/approve.py` | `internal/policy/approve.go` — same grant lifecycle, DB-backed | | `oikos/ledger.py` | `internal/db/` — ledger_entries table + sqlc queries | | `oikos/policy.py` + `policy.yaml` | `internal/policy/` + `seeds/policy.yaml` → DB tables | | `oikos/drift.py` | `internal/scheduler/` — drift detection, writes signals to DB | | `oikos/relations.py` | `internal/ontology/graph.go` — SQL graph traversal | | `oikos/report.py` | `internal/api/routes/` — report endpoints, reads from DB | | `mcp/server.py` | `internal/api/mcp.go` — MCP adapter on top of DB | | `bin/homelab` logic | `internal/api/routes/` — same operations, REST interface | | `oikos/scheduler.py` | `internal/scheduler/` — same probes, goroutines for concurrency | | `oikos/approve.py` Matrix delivery | `internal/notifier/matrix.go` | | *(new)* | `internal/learning/` — feedback, patterns, skills (no Python equivalent) | ## Risks / trade-offs - **Go rewrite** — the existing Python code (~4400 lines) is replaced. The logic and design patterns carry over, but it's a full rewrite. Mitigated by the fact that the Python code is well-documented and the Go structure mirrors it. - **Hermes in Docker** — agent's world is the container. SSH access is the bridge. Hybrid approach (mounted keys now, actuator gateway later). - **PostgreSQL as SPOF** — mitigated by Docker volume persistence + automated `pg_dump` backups (the scheduler can do this once running). - **Learning model cold start** — no patterns/skills exist initially. The agent starts cautious (escalates everything), accumulates feedback, and gradually becomes more autonomous as patterns validate. This is by design — trust is earned. - **Infisical bootstrapping** — SOPS coexists during transition. Keep SOPS as fallback. - **Multi-agent concurrency** — `SELECT ... FOR UPDATE SKIP LOCKED` prevents two actuator passes from acting on the same signal. - **Ontology evolution** — as the homelab changes, entity types and relationship types need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy. ## Audit findings (best practices, security, performance, sanity) A full audit was performed against the plan. Findings are organized by category and severity. **High-severity items must be addressed before implementation; medium items should be addressed during the phase they belong to.** ### Security | ID | Severity | Finding | Recommendation | |---|---|---|---| | S1 | **HIGH** | SSH private keys mounted into Hermes container — a compromised container has unrestricted SSH to all hosts | Use a dedicated, restricted SSH key with `command=` in authorized_keys. Time-box the stopgap. Build the actuator gateway (which brokers SSH per-execution) sooner rather than later. | | S2 | **HIGH** | MCP interface has no auth — any container on the Docker network has full read access to the control plane | Add a shared secret or mTLS between API and Hermes. Bind MCP to a dedicated Docker network, not the default bridge. Never expose the MCP port via Caddy without auth. | | S3 | **HIGH** | Policy DB is mutable via the same API it governs — a compromised API can rewrite its own approval rules (e.g. flip `destructive` to auto-act) | Policy mutations require a meta-approval (dual-control). Add an immutable audit log of policy changes. Startup self-check: alert if policy hash differs from a known-good baseline. | | S4 | **HIGH** | Learning model poisoning — flapping services or misconfigured probes can inject biased feedback to push patterns past the confidence threshold and unlock auto-act for destructive actions | Require human confirmation before pattern transitions to `active`. Cap confidence by sample size (require N≥5 executions). Detect anomalous feedback bursts and quarantine. Never let a skill auto-promote to destructive risk class. | | S5 | **HIGH** | `confirmation_phrase` is a weak auth primitive — if reused, stored in plaintext, or transmitted via Matrix, it's replayable | Make it single-use (one phrase per approval). Store hashed. Transmit only the decision + HMAC, not the phrase. Prefer signed approval tokens. | | S6 | MEDIUM | REST `/exec` and `/decide` endpoints rely on Authentik forward-auth only — if the API port is reachable directly on the mesh, auth disappears | Enforce OIDC JWT validation in the API middleware too (defense in depth). Bind API port to localhost + Caddy only. | | S7 | MEDIUM | No TLS between internal services — Postgres connections are plaintext on the Docker network | TLS to Postgres (server cert verification). Use dedicated Docker networks per trust boundary. Document the threat model: is the Docker network trusted? | | S8 | MEDIUM | Webhook (Gitea → mac-mini deploy) has no auth specified — anyone who can reach the endpoint can trigger arbitrary code execution via `go build` | HMAC signature verification on the webhook. Bind listener to localhost (Gitea reaches via mesh). Run deploy script as non-root. Verify commit signatures before `git pull`. | | S9 | MEDIUM | Infisical bootstrapping has a chicken-and-egg — Infisical's own master key must come from somewhere | Document the bootstrap root of trust explicitly: where the master key lives (mac-mini keychain), how it's backed up, revocation path. Keep SOPS as fallback until Infisical has a tested restore-from-backup drill. | | S10 | MEDIUM | Blast radius of a compromised container is wide — Hermes has SSH + MCP + gateway port + data volume; actuator has SSH + writes to policy/learning tables | Principle of least privilege per container: only the actuator should have SSH, not Hermes. Split networks: data (PG), ops (SSH egress), front (Caddy). Use Docker user namespaces and read-only root filesystems. | ### Performance | ID | Severity | Finding | Recommendation | |---|---|---|---| | P1 | MEDIUM | Recursive CTE `blast_radius` has no cycle guard — cycles (A→B→A) inflate work exponentially with depth | Add a visited-set guard (array accumulator in the recursion). Cap `max_depth` at 3. Add `LIMIT` on the outer query. Consider a closure table for hot-path queries. | | P2 | MEDIUM | Concurrent probes with no concurrency cap — unbounded goroutines could exhaust FDs or hammer slow targets | Bounded worker pool (`errgroup.SetLimit`). Per-target probe timeout. Jitter to avoid thundering herd on shared backends. | | P3 | MEDIUM | Knowledge graph re-ingestion on every deploy is wasteful if no docs changed | Content-hash each doc (store hash on the entity). Skip ingestion if hash unchanged. Run in a single transaction with deferred FK checks. | | P4 | MEDIUM | Pattern extraction frequency unspecified — either too tight (scans whole table) or too loose (patterns lag) | Define cadence (hourly). Use a high-watermark on `feedback.ts`. Add index on `feedback(ts)`. Process only new feedback. | | P5 | **HIGH** | Table growth unaddressed — `state_snapshots` (every entity every 10 min) and `executions`/`feedback` grow indefinitely | Partition `state_snapshots` and `signals` by month. Define TTLs (snapshots > 90 days → aggregate, raw → archive). Add a `prune_*` job in the scheduler. | | P6 | MEDIUM | WebSocket `/events` has no backpressure — a stalled client pins a goroutine and accumulates memory | Bounded channel per subscriber with drop-oldest-on-full. Max subscribers cap. Heartbeat/timeout. Document: best-effort vs. guaranteed delivery. | | P7 | LOW | Deploy script runs `go build` on host AND in Docker — double build, host Go toolchain is a deploy dependency | Standardize on multi-stage Docker build only. Drop host `go build` from deploy script. Keep host toolchain for local dev only. | ### Architecture | ID | Severity | Finding | Recommendation | |---|---|---|---| | A1 | **HIGH** | No testing strategy — 13 workstreams, no mention of unit/integration/property tests | Add a testing workstream: unit tests for `classify.go`, `pattern.go`, `skill.go`; integration tests with testcontainers Postgres; golden-file tests for seed ingest; property test for blast-radius (cycles, depth caps). Coverage gates per package. | | A2 | **HIGH** | No observability — no structured logging, metrics, or traces. Can't inspect why the classifier escalated or how long probes took | At minimum: structured JSON logs (slog) with correlation IDs per execution. A `/metrics` endpoint (even before full Prometheus). `debug=true` flag for full probe payloads. Reconsider Prometheus deferral — "who watches the watcher" requires metrics. | | A3 | **HIGH** | DB backup strategy is a one-liner — the DB is the entire control-plane state (ontology, policy, inventory, learning, ledger) | Define: `pg_dump` cadence (daily + WAL archiving for PITR), off-host storage (push to hubris or object storage, not the same mac-mini volume), encryption at rest, tested restore procedure (monthly drill), retention (30 daily + 12 monthly). This is Phase 1, not "once running." | | A4 | MEDIUM | Seed ingest transactional safety — if API crashes mid-ingest, DB could be in partial state | Run migrations + seed ingest as a distinct init container before the API starts. Wrap each seed file in a single transaction. Add a `seed_version` table to skip already-applied seed versions. | | A5 | MEDIUM | Config management underspecified — unclear what's env vs. DB vs. file vs. Infisical | Define a config hierarchy: defaults → file → env → Infisical (secrets only). Document each service's required config keys. Avoid putting non-secret config in Infisical. | | A6 | MEDIUM | apps/105 fallback has no cutover safety — both stacks can't write to the same state without conflict | Define cutover mode: apps/105 is read-only during coexistence. If rollback needed after new OS has been writing to Postgres, document the reconciliation plan. | | A7 | LOW | Notifier as a separate service is over-engineered initially | Either keep in-process (split later), or document the failure mode: if notifier is down, the API must retry/queue and the operator needs an alternative approval path. | ### Data model | ID | Severity | Finding | Recommendation | |---|---|---|---| | D1 | **HIGH** | `entities.id` as TEXT with manual naming is fragile — renames break the ID; signal IDs require a date-string generator that's race-prone | Use UUID or SERIAL for PK. Keep `name` + `type` as a unique composite for human lookup. Signal IDs: use a DB sequence. | | D2 | MEDIUM | Ontology flexibility sacrifices type safety — `attributes JSONB` has no schema enforcement per entity type | Add a `schema JSONB` column to `entity_types` (JSON Schema). Validate `attributes` against it on insert/update via a trigger or app-layer check. | | D3 | MEDIUM | Entity type evolution has no migration story — renaming/removing an entity type cascades through many FK references | Add `status` (`active`/`deprecated`) to `entity_types`. Forbid hard deletes while instances exist. Provide a `merge` endpoint for renames. | | D4 | MEDIUM | Concurrent writer safety — counters on `patterns` (evidence_count, success_count) have read-modify-write races | Use atomic `UPDATE ... SET evidence_count = evidence_count + 1`. Add `version` optimistic-lock column on `skills` and `patterns`. | | D5 | MEDIUM | Migration rollback not addressed — only one `.down.sql` shown | Commit to forward-only migrations (compensating migrations for rollbacks) and document it, or write and test every `.down.sql`. | | D6 | MEDIUM | Data export for DR is asserted but no endpoint or format defined | Add `GET /api/v1/export` (returns the three YAML files regenerated from DB). Test round-trip: seed → DB → export → seed → DB yields identical state. | | D7 | LOW | `relationships` has no temporal data — can't answer "what depended on authentik last month?" | Add `valid_from`/`valid_to` (nullable = current) for `depends-on` and `hosts` edges. Low priority but cheap to add now. | ### Operational | ID | Severity | Finding | Recommendation | |---|---|---|---| | O1 | **HIGH** | No rollback strategy — if a deploy breaks the API (bad migration, schema bug), recovery is "revert the commit" but the migration may have already run | Define: pre-deploy DB backup snapshot; migration compatibility policy (new code tolerates old schema for one deploy); tested rollback runbook per phase. | | O2 | **HIGH** | "Who watches the watcher" is unresolved — if the OS is down, no alerts fire. No external/orthogonal monitor for the mac-mini or the OS containers | A minimal external watchdog: a cron on apps/105 (or hubris) that curls the API `/healthz` every 5 min and Matrix-pings the operator directly if it fails. This must be *outside* the Docker stack. | | O3 | **HIGH** | No backup/restore runbook — "pg_dump" is mentioned but no procedure, tested restore, or definition of what "restored" means | Write `docs/operations/backup-restore.md`: what's backed up (DB, Infisical, Hermes volume, seed YAMLs), where, how often, how to restore each, quarterly restore drill. | | O4 | **HIGH** | No disaster recovery plan — if the mac-mini dies (disk, theft, water), what's the RTO/RPO? | Define RTO/RPO targets (homelab: RTO 4h, RPO 24h). Name the off-host backup target. Write the DR runbook: fresh mac-mini → install Docker → clone repo → restore Infisical → restore DB → `docker compose up`. | | O5 | MEDIUM | Deploy downtime — `docker compose up -d` recreates containers; API has a brief gap | Use `docker compose up -d --no-deps` for app services. Don't recreate the Postgres container on routine deploys (pin its image). Healthcheck-gated rollout. | | O6 | MEDIUM | Health checks asserted but not specified | Define per-service: API `/healthz` (DB ping), scheduler "last successful probe < 15min ago", notifier "last poll < 60s ago", Hermes "gateway responding". Wire into Docker healthcheck + alerting. | | O7 | MEDIUM | mac-mini single-host failure = total outage — disk failure, macOS update reboot, Docker daemon crash | Automated macOS update deferral/scheduling. Docker `restart: always` on all services. Monitoring heartbeat from external host. Documented cold-start runbook (what comes up first, in what order). | ### Missing | ID | Severity | Finding | Recommendation | |---|---|---|---| | M1 | **HIGH** | No CI/CD — deploys are git-push → webhook → build. No linting, no tests before deploy, no gated merges | Add CI stage (Gitea Actions): `go vet`, `golangci-lint`, `go test ./...`, `docker build` (no push). Gate the webhook on green CI. At minimum, deploy script runs `go test` before `docker compose up`. | | M2 | MEDIUM | No rate limiting on the API — a runaway agent loop or misconfigured skill can hammer the API and DB | Per-caller rate limiting (token bucket) on mutating endpoints. `/exec` needs a per-entity/per-action rate cap to prevent actuator loops. | | M3 | MEDIUM | No access audit for human activity — the ledger records OS actions, but not operator API actions (entity edits, policy changes, approval decisions) | Add an `audit_log` table for all mutating REST calls, including operator identity from OIDC. | | M4 | MEDIUM | No circuit breaker for the actuator — if a target host is unreachable, the actuator keeps attempting SSH and generating failed executions, poisoning the learning model | Circuit breaker per target: after N consecutive failures, back off (exponential) and raise a "target unreachable" signal instead of continuing to execute. | | M5 | MEDIUM | No secret rotation story — SSH keys, Infisical machine tokens, API shared secrets need rotation policies | Define rotation cadences. Document the rotation procedure for each secret class. Add a "secrets expiring" check to the scheduler. | | M6 | LOW | No dependency/supply-chain hygiene — Go modules, Docker base images, Infisical image not pinned or scanned | Pin base images by digest. Run `govulncheck` in CI. Periodically audit `go.sum`. Use distroless or scratch runtime images. | | M7 | LOW | No documented SLOs — no quantitative success criteria (probe latency, API p99, deploy time, alert delivery) | Add a small SLO table: probe interval 10min ±1min, API p99 < 200ms, deploy < 5min, alert delivery < 30s. | ### Top-priority items to address before implementation 1. **S1 + S10** — SSH keys in containers / wide blast radius → build the actuator gateway first, not incrementally 2. **S3 + S4** — Mutable policy DB + learning model poisoning → these compound: a compromised container can rewrite its own rules and inject feedback to unlock auto-act 3. **A1 + A2 + O2** — No tests, no observability, no external watchdog → can't safely run an autonomous agent without all three 4. **A3 + O3 + O4** — Backup/restore/DR is a one-liner for the SPOF Postgres 5. **O1 + M1** — No rollback strategy and no CI gate on deploys ## Verification (end to end) 1. **Ontology:** `seeds/ontology.yaml` ingested — `SELECT * FROM entity_types` shows all types across 3 layers; lifecycle definitions match the state machine diagrams. 2. **DB:** `docker compose up postgres` — all 5 migrations applied; seed ingest populates entities + relationships from `inventory.yaml`. 3. **API:** `curl http://localhost:8090/api/v1/entities?type=service` returns the fleet; MCP `list_services` works via the same endpoint. 4. **Scheduler:** trigger a probe pass — signals in DB, state snapshots written. 5. **Actuator:** raise a test `service-down` signal → actuator classifies → auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated. 6. **Learning:** after N executions of the same (entity_type, action), a pattern appears with confidence score; after enough evidence, a skill is created. 7. **Classifier with learning:** set `autonomy.auto_act: reversible_low` → next similar signal: classifier checks pattern confidence → auto-acts if high, escalates if low. Kill-switch (`auto_act: off`) → always escalates. 8. **Hermes:** connect from another workstation → agent responds, queries API via MCP, can SSH to hubris. 9. **Secrets:** Infisical running, Go services fetch secrets, SOPS files removed. 10. **Deploy:** `git push` → Gitea webhook → `docker compose build + up -d` → changes live, seed ingest syncs any YAML changes to DB. 11. **Knowledge:** MCP `search_knowledge("caddy")` returns docs linked to `entity:service:caddy`. 12. **Cutover:** stop apps/105, verify production traffic only from Docker OS. ## Out of scope (for now) - Oikos Console web UI (deferred — built later on top of the API) - Multi-node deployment (designed for, not implemented) - Vector embeddings / semantic search (structured graph only for now) - SSH-key-signed approval requests - Prometheus / trend signals - Actuator gateway pattern (Phase 2 of hybrid — start with mounted SSH) - Automated skill extraction via LLM (patterns extracted statistically for now; LLM-assisted skill refinement is a future enhancement)