Major revision of the Docker-based homelab OS plan: 1. Go instead of Python — all services rewritten as Go binaries (Gin web framework, sqlc for DB access, goroutines for probes) 2. Ontology-first design — systems modeling with 3 layers: - Infrastructure (physical, compute, network, storage, software) - Governance (identity, secrets, policy) - Cognition (observation, decision, action, knowledge, learning) 7 Mermaid diagrams: layer map, ER diagram, 3 lifecycle state machines, feedback loop, policy model 3. DB-native config — inventory.yaml/ontology.yaml/policy.yaml become seed manifests (bootstrap + DR). The DB is the runtime source of truth, editable via API. Ontology IS the DB schema (entity_types, relationship_types, lifecycle_defs tables). 4. Feedback loop — agent learns from execution: execution → outcome → feedback → pattern → skill → classification Patterns accumulate from execution history, skills codify proven procedures, classifier uses pattern confidence for auto-act decisions. Cold start: agent starts cautious, earns autonomy through evidence. 5. 5 migration groups: ontology meta-schema, entity instances, operations, learning model, policy. Recursive blast_radius SQL function. 6. Phase 0 added: ontology design before any code.
48 KiB
Plan: Oikos — Docker-based agentic homelab OS on mac-mini
Status: Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first design with systems modeling, DB-native config (inventory/ontology/policy as graph metadata, not YAML files), and a feedback loop where the agent learns from execution.
Vision
Convert this repo into a Docker-based agentic homelab OS written in Go. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors:
- Operator (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix.
- Agent — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and learns from every action to improve over time.
All OS services run in Docker containers on mac-mini. The OS is deployed by a git push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path to multi-node later.
Decisions (from operator Q&A, 2026-07-06)
| Question | Decision |
|---|---|
| Repo structure | One repo, reorganize internally. |
| Language | Go — compiled, type-safe, small containers, goroutines for concurrent probes. |
| Hermes runtime | Runs inside Docker as part of the OS stack (gateway mode). |
| Agent → homelab access | Hybrid — mounted SSH keys now, actuator gateway built incrementally. |
| Data storage | PostgreSQL. Inventory/ontology/policy become DB-native graph metadata. |
| Config (inventory/ontology/policy) | DB-native — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend. |
| Knowledge/context | Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data. |
| Operator interface | Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks. |
| Host | mac-mini for now, designed to scale later. |
| Deploy | Git push → Gitea webhook → Docker rebuild + restart. |
bin/homelab CLI |
Replaced by an API. CLI becomes a thin Go client that calls the OS API. |
| Secrets | Migrate from SOPS+age to Infisical. |
| Matrix | Keep for now, abstract the notification layer for future channels. |
| MCP server | Merged into the unified API — one Go service, REST + MCP interfaces. |
| Agent type | Hermes core + custom homelab skills. |
| Feedback loop | Agent learns from execution — outcomes feed back as patterns and skills that inform future decisions. |
| Ontology | Developed first, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles. |
| apps/105 | Keep running as fallback until Docker OS is proven. |
| mac-mini cleanup | Start fresh in a new directory, clean up old artifacts later. |
Ontology — the systems model
The ontology is the foundational layer of the OS. It defines what exists, how things connect, how they change over time, and how the OS learns about them. It is stored in the database as metadata (entity types, relationship types, lifecycle definitions). YAML seed files bootstrap it on first deploy; after that, the DB is authoritative and editable via API.
Design principles (systems modeling)
- Three layers — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's own behavior + learning). Dependencies flow upward: Cognition depends on Governance depends on Infrastructure.
- Everything is an entity — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, changes, skills) are first-class entities, not second-class records.
- Typed relationships with cardinality — edges carry semantics.
hostsis one-to-many;depends-onis many-to-many;documentsis one-to-one. The graph is queryable for blast radius, dependency chains, and knowledge lookup. - Lifecycles are state machines — every entity type has a lifecycle. Infrastructure
entities move through
planned → active → destroyed. Operational entities have their own lifecycles (signals, approvals, patterns, skills). Transitions can require preconditions. - Policy is attached to the ontology — risk classes and approval rules link to entity types and actions. The policy IS part of the model, not a separate file.
- Learning is modeled — executions produce outcomes, outcomes accumulate into patterns, patterns refine skills, skills inform future decisions. This is an explicit, queryable part of the graph.
Layer map
graph TB
subgraph cognition["Layer 3 — Cognition (the OS's behavior + learning)"]
direction LR
OBS["Observation\nsignal, state-snapshot"]
DEC["Decision\nclassification, risk-assessment"]
ACT["Action\nchange, execution, verification"]
GOV["Governance\napproval-request, approval-decision"]
KNOW["Knowledge\ndocument, runbook, lesson"]
LEARN["Learning\npattern, skill, feedback"]
end
subgraph governance["Layer 2 — Governance (who controls what)"]
direction LR
IDENT["Identity\nperson, agent, identity-provider"]
SEC["Secrets\nsecret, key, access-grant"]
POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
end
subgraph infra["Layer 1 — Infrastructure (the managed world)"]
direction LR
PHYS["Physical\nsite, machine, ups, sensor"]
COMP["Compute\nproxmox-host, lxc, vm,\nworkstation, container"]
NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"]
end
cognition -- "observes, acts on, learns about" --> infra
governance -- "governs access to" --> infra
governance -- "constrains" --> cognition
cognition -- "creates + refines" --> governance
Entity relationship diagram — core entities and typed edges
erDiagram
MACHINE ||--o{ PROXMOX_HOST : "is-a"
PROXMOX_HOST ||--o{ LXC : hosts
PROXMOX_HOST ||--o{ VM : hosts
PROXMOX_HOST ||--o{ WORKSTATION : hosts
LXC ||--o{ SERVICE : provides
VM ||--o{ SERVICE : provides
WORKSTATION ||--o{ SERVICE : provides
LXC ||--o{ MOUNT : has
MOUNT }o--|| STORAGE_POOL : "stores-on"
SERVICE ||--o{ INGRESS_ROUTE : "exposed-by"
INGRESS_ROUTE }o--|| CERTIFICATE : "secured-by"
INGRESS_ROUTE }o--|| IDENTITY_PROVIDER : "secured-by"
SERVICE ||--o{ SERVICE : "depends-on"
CONFIG_REPO ||--|| SERVICE : "configured-by"
DEPLOY_PIPELINE ||--|| SERVICE : "deploys-to"
SERVICE ||--o{ SIGNAL : "monitored-by"
SIGNAL ||--o{ EXECUTION : triggers
EXECUTION ||--|| FEEDBACK : produces
FEEDBACK }o--|| PATTERN : "contributes-to"
PATTERN }o--|| SKILL : "informs"
SKILL ||--o{ CLASSIFICATION : "guides"
CLASSIFICATION ||--|| EXECUTION : "precedes"
EXECUTION ||--o| APPROVAL : "requires"
PERSON ||--o{ APPROVAL : "decides"
DOCUMENT ||--o{ SERVICE : "describes"
RUNBOOK ||--o{ SERVICE : "procedure-for"
AGENT ||--o{ EXECUTION : "performs"
PERSON ||--o{ AGENT : "owns"
Infrastructure lifecycle
stateDiagram-v2
[*] --> planned : operator creates entity
planned --> provisioning : IP reserved, storage chosen, doc stub
provisioning --> active : mesh joined, health-check answering, doc complete
active --> migrating : preflight + backup verified
migrating --> active : post-verify, caddy checked, mounts checked
active --> deprecated : replacement live or role retired
deprecated --> destroyed : backups verified, secrets revoked, ingress removed
destroyed --> [*] : archaeology entry recorded
note right of deprecated
Complete only when
zero inbound edges remain
(no depends-on, no routes-to)
end note
Signal lifecycle
stateDiagram-v2
[*] --> raised : scheduler probe or agent finding
raised --> acknowledged : agent or operator sees it
raised --> muted : operator suppresses (TTL)
acknowledged --> acting : actuator starts execution
acting --> resolved : action succeeded, verification passed
acting --> raised : action failed, re-escalated
raised --> resolved : condition cleared (auto-resolve)
muted --> raised : TTL expired
resolved --> [*]
Execution + learning lifecycle
stateDiagram-v2
[*] --> proposed : signal + recommended_action
proposed --> approved : operator approves (if gated)
proposed --> auto_approved : risk class allows auto-act
approved --> executing : actuator runs
auto_approved --> executing : actuator runs
executing --> verified : verification command passed
executing --> failed : execution or verification failed
executing --> timed_out : exceeded duration limit
failed --> rolled_back : rollback procedure executed
verified --> [*] : feedback recorded, pattern updated
failed --> [*] : feedback recorded, pattern updated
rolled_back --> [*] : feedback recorded, pattern updated
timed_out --> [*] : feedback recorded, pattern updated
Feedback / learning model — the cognition loop
flowchart TB
subgraph observe["Observe"]
SIG["Signal raised\n(service down, disk full, drift)"]
end
subgraph decide["Decide"]
CLASS["Classification\nrisk × blast radius × confidence"]
SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"]
CLASS --> SKILL_LOOKUP
end
subgraph act["Act"]
EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"]
VERIFY["Verification\ncheck if action succeeded"]
EXEC --> VERIFY
end
subgraph learn["Learn (feedback loop)"]
OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"]
FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"]
PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"]
SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"]
OUTCOME --> FEEDBACK
FEEDBACK --> PATTERN
PATTERN --> SKILL_REFINE
end
SIG --> CLASS
SKILL_LOOKUP --> EXEC
VERIFY --> OUTCOME
SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP
subgraph escalation["Escalation path"]
APPROVAL["Approval request\n→ Matrix ✅/❌"]
end
CLASS -- "needs approval" --> APPROVAL
APPROVAL -- "approved" --> EXEC
APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"]
Policy model — how risk attaches to entities
flowchart LR
subgraph ontology["Ontology (DB metadata)"]
ET["entity_types\nhost, service, lxc, vm, ..."]
RT["relationship_types\nhosts, provides, depends-on, ..."]
LD["lifecycle_defs\nstates + transitions"]
end
subgraph policy["Policy (DB records)"]
RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"]
RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"]
AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"]
end
subgraph instances["Instances (DB data)"]
ENT["entities\nhost:hubris, service:caddy, ..."]
REL["relationships\nhubris hosts lxc:apps"]
end
ET --> RULES
RC --> RULES
RULES --> ENT
AUTO --> ENT
ET --> ENT
RT --> REL
LD --> ENT
Target architecture
Container stack on mac-mini
graph TB
subgraph mac-mini["mac-mini — Docker host, always-on"]
subgraph services["Docker Compose — Go binaries"]
PG["PostgreSQL 16\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata"]
INF["Infisical\n(secrets manager)"]
HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"]
API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events\n\nPolicy enforcement + risk\nclassification + ledger"]
SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• SSH to hubris/strong"]
NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"]
end
DEPLOY["Gitea webhook →\ndocker compose build + up -d"]
end
HERMES -- MCP --> API
API --> PG
SCHED --> PG
SCHED -- SSH --> HUBRIS
SCHED -- SSH --> STRONG
API -- "escalate" --> NOTIFIER
NOTIFIER -- "alerts + approvals" --> MATRIX
DEPLOY -- "rebuild" --> services
subgraph external["External"]
CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"]
APPS["apps/105 (fallback)"]
HUBRIS["hubris (PVE)"]
STRONG["strong (PVE)"]
GITEA["Gitea (LXC 104)"]
MATRIX["Matrix (LXC 118)"]
WS["Any workstation\nHermes remote → gateway"]
end
CADDY -- reverse_proxy --> API
GITEA -- webhook --> DEPLOY
WS -- "Hermes gateway" --> HERMES
OODA loop — with learning feedback
flowchart LR
OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"]
ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"]
DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"]
DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"]
APPROVE -- "approved" --> ACT
ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"]
LEARN -. "improves confidence" .-> DECIDE
LEARN --> OBSERVE
Deploy flow
flowchart LR
DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"]
GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"]
MACMINI -- "git pull" --> REPO["repo clone"]
MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"]
STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"]
DB-native configuration
The three YAML files — inventory.yaml, ontology.yaml, policy.yaml — become
seed manifests. They bootstrap the DB on first deploy. After that, the DB is the
runtime source of truth, editable via the API. A future frontend can edit all three
directly.
How it works
flowchart TB
subgraph seeds["Seed manifests (git-tracked, YAML)"]
ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"]
INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"]
POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"]
end
subgraph db["PostgreSQL (runtime source of truth)"]
META["entity_types table\nrelationship_types table\nlifecycle_defs table"]
INST["entities table\nrelationships table"]
POLDB["policies table\nrisk_classes table\nautonomy_settings table"]
end
INGEST["Seed ingest (on deploy)\nidempotent upsert"]
ONTO_YAML --> INGEST --> META
INV_YAML --> INGEST --> INST
POL_YAML --> INGEST --> POLDB
API_EDIT["API edits\n(POST/PUT/PATCH)"]
API_EDIT --> META
API_EDIT --> INST
API_EDIT --> POLDB
EXPORT["Export to YAML\n(for DR / version control)"]
META --> EXPORT
INST --> EXPORT
POLDB --> EXPORT
Why DB-native
- Querying — the agent can ask "what services depend on authentik?" as a graph query, not a YAML parse. Blast-radius walks are SQL, not file reads.
- Mutation — adding a service, updating a lifecycle state, changing a policy rule are DB transactions with audit trail, not file edits + git commits.
- Consistency — the ontology, inventory, and policy are always in sync (same DB, same transaction). No drift between what the YAML says and what the runtime sees.
- Future frontend — a UI can edit entities, relationships, and policies directly via the API. No need to generate/edit YAML files.
- Version control — seed YAML files are still git-tracked for bootstrap and DR. The API can export the current DB state back to YAML for commit.
Repo layout (Go project)
/ # repo root
├── docker-compose.yml # the OS stack definition
├── Makefile # build, test, deploy targets
├── go.mod # Go module definition
├── go.sum
├── cmd/ # binary entrypoints (one per service)
│ ├── api/ # Oikos API server
│ │ └── main.go
│ ├── scheduler/ # Observe + Act loop
│ │ └── main.go
│ └── notifier/ # Notification service
│ └── main.go
├── internal/ # private packages (not importable)
│ ├── db/ # database layer
│ │ ├── queries/ # sqlc SQL queries
│ │ ├── models.go # generated Go types
│ │ └── db.go # connection pool, migrations
│ ├── ontology/ # ontology types + meta-schema
│ │ ├── types.go # EntityType, RelationshipType, LifecycleDef
│ │ ├── graph.go # graph traversal (blast radius, dependencies)
│ │ └── ingest.go # YAML seed → DB ingest
│ ├── api/ # HTTP + MCP server
│ │ ├── server.go # Gin app setup
│ │ ├── routes/ # REST handlers
│ │ │ ├── hosts.go
│ │ │ ├── services.go
│ │ │ ├── signals.go
│ │ │ ├── approvals.go
│ │ │ ├── exec.go
│ │ │ └── knowledge.go
│ │ └── mcp.go # MCP protocol adapter (JSON-RPC over SSE)
│ ├── policy/ # risk classification + approval
│ │ ├── classify.go # risk × blast × confidence
│ │ ├── approve.go # approval request + grant lifecycle
│ │ └── autonomy.go # kill-switch, never-auto-act list
│ ├── scheduler/ # Observe stage
│ │ ├── probe.go # HTTP health, disk, drift
│ │ └── signal.go # raise/resolve signals in DB
│ ├── actuator/ # Act stage
│ │ ├── act.go # read signals, classify, execute or escalate
│ │ ├── execute.go # SSH execution + verification
│ │ └── guard.go # loop-guard, retry caps
│ ├── learning/ # feedback loop (the learning model)
│ │ ├── feedback.go # record outcome + lesson from execution
│ │ ├── pattern.go # extract/validate patterns from feedback
│ │ └── skill.go # create/refine skills from patterns
│ ├── notifier/ # notification abstraction
│ │ ├── notifier.go # interface
│ │ └── matrix.go # Matrix implementation
│ └── config/ # config loading (env, files)
│ └── config.go
├── migrations/ # SQL migrations (golang-migrate format)
│ ├── 001_ontology.up.sql # meta-schema (entity_types, relationship_types, ...)
│ ├── 001_ontology.down.sql
│ ├── 002_instances.up.sql # entities, relationships
│ ├── 003_operations.up.sql # signals, approvals, executions, feedback
│ ├── 004_learning.up.sql # patterns, skills
│ └── 005_policy.up.sql # policies, risk_classes, autonomy
├── seeds/ # YAML seed manifests (bootstrap + DR)
│ ├── ontology.yaml # entity types, relationship types, lifecycles
│ ├── inventory.yaml # entity instances (hosts, services, etc.)
│ └── policy.yaml # risk classes, approval rules, autonomy
├── docs/ # narrative docs (ingested into knowledge graph)
│ ├── containers/
│ ├── hosts/
│ ├── infrastructure/
│ └── investigations/
├── hermes/ # Hermes agent config + skills
│ ├── config.yaml
│ ├── SOUL.md
│ └── skills/
│ └── homelab-ops/
│ └── SKILL.md
├── compose/ # Docker build contexts
│ ├── api/Dockerfile
│ ├── scheduler/Dockerfile
│ ├── hermes/Dockerfile
│ └── postgres/init.sql
└── scripts/ # utility scripts
├── migrate-sops.sh # one-time SOPS → Infisical migration
└── import-legacy.sh # import existing signals/ledger JSONL
Database schema
The schema is the ontology made concrete. Five migration groups, each adding a layer.
Migration 1: Ontology meta-schema
-- The meta-graph: defines what entity types and relationship types can exist.
-- This IS the ontology, stored in the DB, editable via API.
CREATE TABLE entity_types (
name TEXT PRIMARY KEY, -- 'host', 'service', 'signal', 'pattern'
domain TEXT NOT NULL, -- 'physical', 'compute', 'network', ...
layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition'
description TEXT,
lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle)
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE TABLE relationship_types (
name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on'
inverse TEXT, -- 'runs-on', 'provided-by'
source_type TEXT REFERENCES entity_types(name),
target_type TEXT REFERENCES entity_types(name),
cardinality TEXT NOT NULL, -- 'one-to-one', 'one-to-many', 'many-to-many'
description TEXT,
created_at TIMESTAMPTZ DEFAULT now()
);
CREATE TABLE lifecycle_defs (
id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', 'execution', ...
states TEXT[] NOT NULL, -- ordered states
default_state TEXT NOT NULL,
transitions JSONB NOT NULL, -- {"from": {"to": {"requires": [...]}}}
created_at TIMESTAMPTZ DEFAULT now()
);
Migration 2: Entity instances (the inventory graph)
-- Entity instances — the actual hosts, services, signals, patterns, etc.
-- This replaces inventory.yaml as the runtime source of truth.
CREATE TABLE entities (
id TEXT PRIMARY KEY, -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001'
type TEXT NOT NULL REFERENCES entity_types(name),
name TEXT NOT NULL, -- 'hubris', 'caddy', 'disk-threshold'
state TEXT, -- lifecycle state (e.g. 'active', 'raised')
attributes JSONB NOT NULL DEFAULT '{}', -- type-specific data (IP, mesh addr, port, ...)
parent_id TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host)
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_entities_type ON entities(type);
CREATE INDEX idx_entities_state ON entities(state);
CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes);
-- Relationship instances — the typed edges of the graph
CREATE TABLE relationships (
source_id TEXT NOT NULL REFERENCES entities(id),
target_id TEXT NOT NULL REFERENCES entities(id),
type TEXT NOT NULL REFERENCES relationship_types(name),
attributes JSONB,
created_at TIMESTAMPTZ DEFAULT now(),
PRIMARY KEY (source_id, target_id, type)
);
CREATE INDEX idx_rel_source ON relationships(source_id);
CREATE INDEX idx_rel_target ON relationships(target_id);
CREATE INDEX idx_rel_type ON relationships(type);
-- Recursive graph traversal function (blast radius, dependency chains)
CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3)
RETURNS TABLE(entity_id TEXT, depth INT) AS $$
WITH RECURSIVE walk AS (
SELECT start_id::TEXT AS entity_id, 0::INT AS depth
UNION
SELECT r.target_id::TEXT, w.depth + 1
FROM relationships r
JOIN walk w ON r.source_id = w.entity_id
WHERE w.depth < max_depth
)
SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
$$ LANGUAGE sql STABLE;
Migration 3: Operations (signals, approvals, ledger, state)
-- Signals — now entities in the graph, with a dedicated table for indexed querying
-- (the entity row is the canonical record; this table is a fast lookup)
CREATE TABLE signals (
entity_id TEXT PRIMARY KEY REFERENCES entities(id),
kind TEXT NOT NULL,
severity TEXT NOT NULL, -- info, warning, critical
target_entity_id TEXT REFERENCES entities(id), -- the infrastructure entity this is about
evidence TEXT,
likely_cause TEXT,
recommended_action JSONB,
verification TEXT,
state TEXT NOT NULL DEFAULT 'raised',
mute_until TIMESTAMPTZ,
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_signals_state ON signals(state);
CREATE INDEX idx_signals_target ON signals(target_entity_id);
CREATE INDEX idx_signals_severity ON signals(severity);
-- Approvals
CREATE TABLE approvals (
id TEXT PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id), -- entity to act on
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'pending', -- pending, approved, denied, expired
ttl INTERVAL NOT NULL DEFAULT '1 hour',
decided_at TIMESTAMPTZ,
decided_by TEXT REFERENCES entities(id), -- person entity
confirmation_phrase TEXT
);
-- Change ledger (high-level record, links to execution for detail)
CREATE TABLE ledger_entries (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id),
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
result TEXT, -- ok, failed, escalated
approval_id TEXT REFERENCES approvals(id),
execution_id INTEGER, -- FK to executions (migration 4)
agent_id TEXT REFERENCES entities(id),
notes TEXT
);
-- State snapshots (replaces oikos/state.json)
CREATE TABLE state_snapshots (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id),
health TEXT, -- healthy, degraded, down, unknown
data JSONB
);
Migration 4: Learning model (executions, feedback, patterns, skills)
-- Executions — detailed record of each action the OS performs
CREATE TABLE executions (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
signal_entity_id TEXT REFERENCES entities(id), -- signal that triggered this
target_entity_id TEXT REFERENCES entities(id), -- entity acted upon
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
approval_id TEXT REFERENCES approvals(id),
agent_id TEXT REFERENCES entities(id), -- who/what executed
skill_id TEXT REFERENCES entities(id), -- skill used (if any)
status TEXT NOT NULL DEFAULT 'queued', -- queued, running, completed, failed, timed-out
result JSONB, -- detailed result data
duration_ms INTEGER,
verified BOOLEAN DEFAULT false,
started_at TIMESTAMPTZ,
completed_at TIMESTAMPTZ
);
CREATE INDEX idx_exec_target ON executions(target_entity_id);
CREATE INDEX idx_exec_status ON executions(status);
CREATE INDEX idx_exec_action ON executions(action);
-- Feedback — what was learned from an execution
CREATE TABLE feedback (
id SERIAL PRIMARY KEY,
execution_id INTEGER REFERENCES executions(id),
ts TIMESTAMPTZ DEFAULT now(),
outcome TEXT NOT NULL, -- success, failure, partial, unexpected
observation TEXT, -- what happened vs what was expected
lesson TEXT, -- extractable lesson
unexpected_side_effects TEXT[],
tags TEXT[]
);
CREATE INDEX idx_feedback_execution ON feedback(execution_id);
CREATE INDEX idx_feedback_outcome ON feedback(outcome);
-- Patterns — generalized rules extracted from accumulated feedback
CREATE TABLE patterns (
id TEXT PRIMARY KEY, -- 'pat-2026-07-06-001'
ts TIMESTAMPTZ DEFAULT now(),
entity_type TEXT REFERENCES entity_types(name), -- applies to this type
action TEXT NOT NULL, -- 'restart', 'deploy', etc.
pattern TEXT NOT NULL, -- 'service X recovers within 30s after restart'
confidence REAL DEFAULT 0.5, -- 0.0 to 1.0
evidence_count INTEGER DEFAULT 1, -- how many executions support this
success_count INTEGER DEFAULT 0,
failure_count INTEGER DEFAULT 0,
status TEXT DEFAULT 'hypothesized', -- hypothesized, validated, active, deprecated
last_validated_at TIMESTAMPTZ
);
CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action);
CREATE INDEX idx_patterns_status ON patterns(status);
-- Skills — codified procedures refined through feedback
CREATE TABLE skills (
id TEXT PRIMARY KEY, -- 'skill-restart-service', 'skill-deploy-lxc'
name TEXT NOT NULL,
ts TIMESTAMPTZ DEFAULT now(),
procedure TEXT NOT NULL, -- the codified steps (markdown or structured)
applies_to TEXT REFERENCES entity_types(name),
pattern_ids TEXT[], -- patterns that inform this skill
status TEXT DEFAULT 'drafted', -- drafted, tested, active, refined, deprecated
version INTEGER DEFAULT 1,
success_rate REAL, -- rolling success rate
last_used_at TIMESTAMPTZ
);
CREATE INDEX idx_skills_type ON skills(applies_to);
CREATE INDEX idx_skills_status ON skills(status);
Migration 5: Policy (DB-native risk + approval rules)
-- Risk classes — the four-level safety model
CREATE TABLE risk_classes (
name TEXT PRIMARY KEY, -- 'read_only', 'reversible_low', etc.
description TEXT,
approval_required TEXT NOT NULL DEFAULT 'none', -- none, operator, operator_confirmed
ledger BOOLEAN DEFAULT false,
autonomy_allowed BOOLEAN DEFAULT false -- can agent auto-act at this risk level?
);
-- Approval rules — entity_type + action → risk_class + requirements
CREATE TABLE approval_rules (
id SERIAL PRIMARY KEY,
entity_type TEXT REFERENCES entity_types(name), -- applies to this entity type
action TEXT NOT NULL, -- 'restart', 'deploy', 'destroy'
risk_class TEXT NOT NULL REFERENCES risk_classes(name),
autonomy_level TEXT NOT NULL DEFAULT 'auto', -- 'auto', 'escalate', 'never'
scope_entity TEXT REFERENCES entities(id), -- optional: specific entity only
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now(),
UNIQUE(entity_type, action)
);
-- Autonomy settings — global kill-switch + per-entity overrides
CREATE TABLE autonomy_settings (
key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.caddy'
value TEXT NOT NULL, -- 'off', 'reversible_low', 'true', 'false'
updated_at TIMESTAMPTZ DEFAULT now()
);
Workstreams
1. Ontology definition + seed manifests (seeds/, internal/ontology/)
Before any code, finalize the ontology. The current ontology.yaml has 8 domains
and 14 relationship types. The new ontology adds:
- Layer 3 entities (cognition): signal, execution, feedback, pattern, skill, approval, classification, document, runbook
- New relationships:
triggers,produces,contributes-to,informs,guides,precedes,performs,procedure-for,learned-from - Lifecycle definitions for operational entities (signals, executions, patterns, skills) — not just infrastructure
Deliverables:
seeds/ontology.yaml— entity types, relationship types, lifecycle definitions (seeded intoentity_types,relationship_types,lifecycle_defson deploy)seeds/inventory.yaml— adapted from currentinventory.yaml(seeded intoentities+relationships)seeds/policy.yaml— adapted from currentpolicy.yaml(seeded intorisk_classes,approval_rules,autonomy_settings)internal/ontology/ingest.go— idempotent seed → DB ingest
2. Database layer + migrations (migrations/, internal/db/)
- Write the 5 migrations above
- Set up sqlc for type-safe Go database access
- Connection pool, migration runner
- Graph traversal queries (blast radius, dependency chains, knowledge lookup)
- Import script for existing
signals/*.jsonl,ledger/*.jsonl→ DB
3. Unified API server — Go (cmd/api/, internal/api/)
One Go binary (Gin web framework) exposing REST + MCP from the same codebase.
MCP interface (internal/api/mcp.go):
- JSON-RPC over SSE, compatible with Hermes MCP client
- Tools:
get_host,list_services,search_knowledge,get_entity,get_relations,get_blast_radius,get_signal_history,get_ledger,get_patterns,get_skills,get_state_snapshot - All read from PostgreSQL
REST interface (internal/api/routes/):
GET /api/v1/entities— list entities (filter by type, state, domain)GET /api/v1/entities/{id}— entity detail + relationshipsPOST /api/v1/entities— create entity (creates inventory entry)PATCH /api/v1/entities/{id}— update entity (state transition, attributes)GET /api/v1/signals— list signals (filter by state, severity, entity)POST /api/v1/signals/{id}/ack— acknowledgePOST /api/v1/signals/{id}/resolve— resolveGET /api/v1/approvals— pending approvalsPOST /api/v1/approvals/{id}/decide— approve/deny (Authentik-gated)POST /api/v1/exec— gated execution (classify → check approval → execute → feedback)GET /api/v1/patterns— list patterns (filter by entity_type, action, status)GET /api/v1/skills— list skills (filter by applies_to, status)GET /api/v1/knowledge/{entity_id}— knowledge graph queryGET /api/v1/knowledge/search?q=...— search knowledge graphGET /api/v1/ontology— list entity types, relationship types, lifecyclesPOST /api/v1/ontology/entity_types— create entity type (extend the schema)WS /api/v1/events— real-time stream (signals, approvals, executions, feedback)
Policy enforcement (internal/policy/classify.go):
- Every mutating endpoint classifies the action via the policy DB
- Risk class → approval check → autonomy check
- All mutations write to the ledger automatically
Auth:
- MCP interface: no auth (internal, container-to-container)
- REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints
- Internal: shared secret (Docker network)
4. Scheduler + Actuator — Go (cmd/scheduler/, internal/scheduler/, internal/actuator/)
Scheduler (Observe) — Go service with goroutines for concurrent probes:
- HTTP health probes (concurrent, with timeouts)
- Disk usage probes (SSH to hubris/strong)
- Drift detection (inventory vs live state)
- Writes signals + state snapshots to DB
- Runs on a 10-min ticker
Actuator (Act) — Go service, the control loop:
- Reads open signals with
recommended_action - For each: classify via
internal/policy/classify.go- auto-act: look up skill for (entity_type, action) → follow procedure → execute via SSH → verify → record execution → generate feedback → update patterns
- escalate: create approval request → notify via Matrix → acknowledge signal
- Loop-guard: check execution history per (entity, action) to cap auto-retries
(
SELECT ... FOR UPDATE SKIP LOCKEDfor concurrency safety) - Autonomy kill-switch: check
autonomy_settingstable
5. Learning engine — Go (internal/learning/)
The feedback loop that makes the agent improve over time.
Feedback recording (internal/learning/feedback.go):
- After every execution, evaluate the outcome:
- Did the verification command pass? → success
- Did it fail? → failure
- Did it partially work? → partial
- Did something unexpected happen? → unexpected
- Record a feedback entry with: outcome, observation (what happened vs expected), lesson (extractable insight), unexpected_side_effects
Pattern extraction (internal/learning/pattern.go):
- Periodically scan accumulated feedback for (entity_type, action) pairs
- When N+ executions share a similar outcome, extract a pattern:
- "Restarting service:X typically takes 15s and succeeds"
- "Deploying to LXC:Y via webhook has 30% failure rate, retry helps"
- Patterns start as
hypothesized, move tovalidatedafter enough evidence, thenactive(used by the decision classifier) - Confidence score = success_count / evidence_count, adjusted by recency
Skill management (internal/learning/skill.go):
- When a pattern reaches
activestatus with confidence > 0.7, create or refine a skill for that (entity_type, action) pair - Skills codify the best-known procedure (what steps to take, what to verify, expected duration, known failure modes)
- Skills are versioned — each refinement increments the version
- The actuator looks up skills before executing: if a skill exists, follow it; if not, use the default procedure and generate feedback for future pattern extraction
How the classifier uses learning (internal/policy/classify.go):
- Confidence scoring now checks patterns + skills, not just raw ledger history:
- If a pattern exists for (entity_type, action) with high confidence → boost auto-act confidence
- If patterns show frequent failures → lower confidence, escalate
- If a skill exists → higher confidence (proven procedure available)
- This is the closed loop: execution → feedback → pattern → skill → classification → execution (better informed each time)
6. Knowledge graph ingestion (internal/ontology/ingest.go)
After the ontology is defined and the DB schema is in place:
- On deploy, walk
docs/directory - Parse each markdown file:
- Extract frontmatter for metadata (entity type, tags, relations)
- Infer entity relationships from path conventions:
docs/containers/105-apps.md→ relationship toentity:lxc:apps - Extract cross-references (markdown links) → relationships
- Create knowledge entities in the
entitiestable (type =document,runbook,investigation, etc.) withdocuments/procedure-foredges to infrastructure entities - Idempotent — safe to re-run on every deploy
Agent access:
- MCP tool
search_knowledge(query)— full-text search on knowledge entities - MCP tool
get_entity_knowledge(entity_id)— all docs related to an entity - MCP tool
get_relations(entity_id)— graph traversal (blast radius, dependencies)
7. Hermes agent container (compose/hermes/)
- Hermes Agent runtime in a Docker container, gateway mode
- Config:
hermes/config.yaml(providers, models, gateway port) - Persona:
hermes/SOUL.md(homelab-specific) - Skills:
hermes/skills/homelab-ops/SKILL.md— how to use the Oikos API, classify actions, request approvals, query the knowledge graph - Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid)
- Gateway port 8092 — workstations connect remotely
- Hermes data volume for persistent state
8. Infisical secrets migration
Same as rev 1:
- Stand up Infisical in the Docker stack
- Migrate SOPS secrets (one-time decrypt + import)
- Wire all Go services to Infisical via machine identity
- Retire SOPS + age keys
9. Notifier — Go (cmd/notifier/, internal/notifier/)
Interface (Go):
type Notifier interface {
SendAlert(ctx context.Context, signal Signal) error
SendApprovalRequest(ctx context.Context, approval Approval) error
ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error)
}
Matrix implementation:
- Sends alerts to
@dtoro:avisperovia Synapse (LXC 118) - Approval requests as messages with ✅/❌ reactions
- Listens for reactions to record decisions
- Pluggable — future implementations (webhook, email) register via config
10. Docker build + deploy pipeline
- Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch)
docker compose buildfrom repo root- Gitea webhook on push to
main→ deploy script on mac-mini - Deploy script:
git pull && go build ./... && docker compose build && docker compose up -d - Seed ingest runs as part of the API startup (idempotent)
- Health checks on each service
11. Ingress re-point (Caddy)
dtoro/caddy-conf: pointmcp.hubris.network+oikos.hubris.network→ mac-mini mesh IP :8090 (API)- Future:
hermes.hubris.network→ mac-mini:8092 - DNS and public URLs unchanged
12. mac-mini host setup
- New directory:
~/oikos-os/— repo clone +docker composeworking dir - Existing
/opt/homelab-context/stays untouched until cleanup phase - Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys
- Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone
13. Decommission apps/105 (deferred)
Keep apps/105 running as fallback. Cutover checklist when ready:
- Verify Docker OS serves all traffic
systemctl disable --nowOikos services on apps/105- Remove old checkouts
- Update Caddy backends exclusively to mac-mini
- Remove/retarget Gitea webhooks
Phasing
Phase 0 — Ontology design (no code):
- Finalize entity types, relationship types, lifecycles
- Write seed manifests (
seeds/ontology.yaml,seeds/inventory.yaml,seeds/policy.yaml) - Review diagrams with operator
Phase 1 — Foundation (Go + DB):
- Go module setup, project structure
- PostgreSQL migrations (all 5)
- Seed ingest pipeline (YAML → DB)
- sqlc queries for core operations
- Import existing signals/ledger data
Phase 2 — API (Go):
- Gin server with REST routes
- MCP protocol adapter
- Policy enforcement middleware
- Knowledge graph ingestion from docs/
Phase 3 — Control loop (Go):
- Scheduler (Observe) — probes, signals, state snapshots
- Actuator (Act) — classify, execute, verify
- Learning engine — feedback, patterns, skills
Phase 4 — Agent (Hermes container):
- Hermes Docker image, gateway config
- Homelab skills
- Connect from workstation, verify MCP + SSH
Phase 5 — Secrets (Infisical):
- Stand up Infisical, migrate SOPS, wire services
Phase 6 — Deploy + cutover:
- Docker Compose, Gitea webhook, Caddy re-point
- End-to-end verification
- Stop apps/105, clean up mac-mini
Reuse (logic carried over, rewritten in Go)
| Existing Python | What it becomes in Go |
|---|---|
oikos/decide.py |
internal/policy/classify.go — same scoring logic, reads from DB + patterns |
oikos/signal.py |
internal/scheduler/signal.go — same lifecycle, DB-backed |
oikos/approve.py |
internal/policy/approve.go — same grant lifecycle, DB-backed |
oikos/ledger.py |
internal/db/ — ledger_entries table + sqlc queries |
oikos/policy.py + policy.yaml |
internal/policy/ + seeds/policy.yaml → DB tables |
oikos/drift.py |
internal/scheduler/ — drift detection, writes signals to DB |
oikos/relations.py |
internal/ontology/graph.go — SQL graph traversal |
oikos/report.py |
internal/api/routes/ — report endpoints, reads from DB |
mcp/server.py |
internal/api/mcp.go — MCP adapter on top of DB |
bin/homelab logic |
internal/api/routes/ — same operations, REST interface |
oikos/scheduler.py |
internal/scheduler/ — same probes, goroutines for concurrency |
oikos/approve.py Matrix delivery |
internal/notifier/matrix.go |
| (new) | internal/learning/ — feedback, patterns, skills (no Python equivalent) |
Risks / trade-offs
- Go rewrite — the existing Python code (~4400 lines) is replaced. The logic and design patterns carry over, but it's a full rewrite. Mitigated by the fact that the Python code is well-documented and the Go structure mirrors it.
- Hermes in Docker — agent's world is the container. SSH access is the bridge. Hybrid approach (mounted keys now, actuator gateway later).
- PostgreSQL as SPOF — mitigated by Docker volume persistence + automated
pg_dumpbackups (the scheduler can do this once running). - Learning model cold start — no patterns/skills exist initially. The agent starts cautious (escalates everything), accumulates feedback, and gradually becomes more autonomous as patterns validate. This is by design — trust is earned.
- Infisical bootstrapping — SOPS coexists during transition. Keep SOPS as fallback.
- Multi-agent concurrency —
SELECT ... FOR UPDATE SKIP LOCKEDprevents two actuator passes from acting on the same signal. - Ontology evolution — as the homelab changes, entity types and relationship types need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy.
Verification (end to end)
- Ontology:
seeds/ontology.yamlingested —SELECT * FROM entity_typesshows all types across 3 layers; lifecycle definitions match the state machine diagrams. - DB:
docker compose up postgres— all 5 migrations applied; seed ingest populates entities + relationships frominventory.yaml. - API:
curl http://localhost:8090/api/v1/entities?type=servicereturns the fleet; MCPlist_servicesworks via the same endpoint. - Scheduler: trigger a probe pass — signals in DB, state snapshots written.
- Actuator: raise a test
service-downsignal → actuator classifies → auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated. - Learning: after N executions of the same (entity_type, action), a pattern appears with confidence score; after enough evidence, a skill is created.
- Classifier with learning: set
autonomy.auto_act: reversible_low→ next similar signal: classifier checks pattern confidence → auto-acts if high, escalates if low. Kill-switch (auto_act: off) → always escalates. - Hermes: connect from another workstation → agent responds, queries API via MCP, can SSH to hubris.
- Secrets: Infisical running, Go services fetch secrets, SOPS files removed.
- Deploy:
git push→ Gitea webhook →docker compose build + up -d→ changes live, seed ingest syncs any YAML changes to DB. - Knowledge: MCP
search_knowledge("caddy")returns docs linked toentity:service:caddy. - Cutover: stop apps/105, verify production traffic only from Docker OS.
Out of scope (for now)
- Oikos Console web UI (deferred — built later on top of the API)
- Multi-node deployment (designed for, not implemented)
- Vector embeddings / semantic search (structured graph only for now)
- SSH-key-signed approval requests
- Prometheus / trend signals
- Actuator gateway pattern (Phase 2 of hybrid — start with mounted SSH)
- Automated skill extraction via LLM (patterns extracted statistically for now; LLM-assisted skill refinement is a future enhancement)