Major revision of the Docker-based homelab OS plan: 1. Go instead of Python — all services rewritten as Go binaries (Gin web framework, sqlc for DB access, goroutines for probes) 2. Ontology-first design — systems modeling with 3 layers: - Infrastructure (physical, compute, network, storage, software) - Governance (identity, secrets, policy) - Cognition (observation, decision, action, knowledge, learning) 7 Mermaid diagrams: layer map, ER diagram, 3 lifecycle state machines, feedback loop, policy model 3. DB-native config — inventory.yaml/ontology.yaml/policy.yaml become seed manifests (bootstrap + DR). The DB is the runtime source of truth, editable via API. Ontology IS the DB schema (entity_types, relationship_types, lifecycle_defs tables). 4. Feedback loop — agent learns from execution: execution → outcome → feedback → pattern → skill → classification Patterns accumulate from execution history, skills codify proven procedures, classifier uses pattern confidence for auto-act decisions. Cold start: agent starts cautious, earns autonomy through evidence. 5. 5 migration groups: ontology meta-schema, entity instances, operations, learning model, policy. Recursive blast_radius SQL function. 6. Phase 0 added: ontology design before any code.
1071 lines
48 KiB
Markdown
1071 lines
48 KiB
Markdown
# Plan: Oikos — Docker-based agentic homelab OS on mac-mini
|
||
|
||
**Status:** Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the
|
||
Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first
|
||
design with systems modeling, DB-native config (inventory/ontology/policy as graph
|
||
metadata, not YAML files), and a feedback loop where the agent learns from execution.
|
||
|
||
## Vision
|
||
|
||
Convert this repo into a **Docker-based agentic homelab OS** written in **Go**. The OS
|
||
is a set of containerized services that manage the homelab autonomously, with the
|
||
operator in control. Two actors:
|
||
|
||
- **Operator** (dtoro) — owns the homelab, expresses intent ("install X", "restart
|
||
Y"), approves destructive actions. Connects from any workstation via remote Hermes
|
||
or Matrix.
|
||
- **Agent** — Hermes core + custom homelab skills, running in Docker. Executes orders,
|
||
monitors the lab, escalates when unsure, and **learns from every action** to improve
|
||
over time.
|
||
|
||
All OS services run in Docker containers on mac-mini. The OS is deployed by a git
|
||
push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path
|
||
to multi-node later.
|
||
|
||
## Decisions (from operator Q&A, 2026-07-06)
|
||
|
||
| Question | Decision |
|
||
|---|---|
|
||
| Repo structure | One repo, reorganize internally. |
|
||
| Language | **Go** — compiled, type-safe, small containers, goroutines for concurrent probes. |
|
||
| Hermes runtime | Runs inside Docker as part of the OS stack (gateway mode). |
|
||
| Agent → homelab access | Hybrid — mounted SSH keys now, actuator gateway built incrementally. |
|
||
| Data storage | PostgreSQL. Inventory/ontology/policy become DB-native graph metadata. |
|
||
| Config (inventory/ontology/policy) | **DB-native** — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend. |
|
||
| Knowledge/context | Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data. |
|
||
| Operator interface | Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks. |
|
||
| Host | mac-mini for now, designed to scale later. |
|
||
| Deploy | Git push → Gitea webhook → Docker rebuild + restart. |
|
||
| `bin/homelab` CLI | Replaced by an API. CLI becomes a thin Go client that calls the OS API. |
|
||
| Secrets | Migrate from SOPS+age to Infisical. |
|
||
| Matrix | Keep for now, abstract the notification layer for future channels. |
|
||
| MCP server | Merged into the unified API — one Go service, REST + MCP interfaces. |
|
||
| Agent type | Hermes core + custom homelab skills. |
|
||
| Feedback loop | **Agent learns from execution** — outcomes feed back as patterns and skills that inform future decisions. |
|
||
| Ontology | **Developed first**, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles. |
|
||
| apps/105 | Keep running as fallback until Docker OS is proven. |
|
||
| mac-mini cleanup | Start fresh in a new directory, clean up old artifacts later. |
|
||
|
||
## Ontology — the systems model
|
||
|
||
The ontology is the foundational layer of the OS. It defines what exists, how things
|
||
connect, how they change over time, and how the OS learns about them. It is stored
|
||
**in the database** as metadata (entity types, relationship types, lifecycle
|
||
definitions). YAML seed files bootstrap it on first deploy; after that, the DB is
|
||
authoritative and editable via API.
|
||
|
||
### Design principles (systems modeling)
|
||
|
||
1. **Three layers** — Infrastructure (the managed world), Governance (who controls
|
||
what), Cognition (the OS's own behavior + learning). Dependencies flow upward:
|
||
Cognition depends on Governance depends on Infrastructure.
|
||
2. **Everything is an entity** — if it can break, be changed, or hold data, it has an
|
||
entity type and edges. The OS's own objects (signals, changes, skills) are
|
||
first-class entities, not second-class records.
|
||
3. **Typed relationships with cardinality** — edges carry semantics. `hosts` is
|
||
one-to-many; `depends-on` is many-to-many; `documents` is one-to-one. The graph
|
||
is queryable for blast radius, dependency chains, and knowledge lookup.
|
||
4. **Lifecycles are state machines** — every entity type has a lifecycle. Infrastructure
|
||
entities move through `planned → active → destroyed`. Operational entities have
|
||
their own lifecycles (signals, approvals, patterns, skills). Transitions can require
|
||
preconditions.
|
||
5. **Policy is attached to the ontology** — risk classes and approval rules link to
|
||
entity types and actions. The policy IS part of the model, not a separate file.
|
||
6. **Learning is modeled** — executions produce outcomes, outcomes accumulate into
|
||
patterns, patterns refine skills, skills inform future decisions. This is an
|
||
explicit, queryable part of the graph.
|
||
|
||
### Layer map
|
||
|
||
```mermaid
|
||
graph TB
|
||
subgraph cognition["Layer 3 — Cognition (the OS's behavior + learning)"]
|
||
direction LR
|
||
OBS["Observation\nsignal, state-snapshot"]
|
||
DEC["Decision\nclassification, risk-assessment"]
|
||
ACT["Action\nchange, execution, verification"]
|
||
GOV["Governance\napproval-request, approval-decision"]
|
||
KNOW["Knowledge\ndocument, runbook, lesson"]
|
||
LEARN["Learning\npattern, skill, feedback"]
|
||
end
|
||
|
||
subgraph governance["Layer 2 — Governance (who controls what)"]
|
||
direction LR
|
||
IDENT["Identity\nperson, agent, identity-provider"]
|
||
SEC["Secrets\nsecret, key, access-grant"]
|
||
POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
|
||
end
|
||
|
||
subgraph infra["Layer 1 — Infrastructure (the managed world)"]
|
||
direction LR
|
||
PHYS["Physical\nsite, machine, ups, sensor"]
|
||
COMP["Compute\nproxmox-host, lxc, vm,\nworkstation, container"]
|
||
NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
|
||
STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
|
||
SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"]
|
||
end
|
||
|
||
cognition -- "observes, acts on, learns about" --> infra
|
||
governance -- "governs access to" --> infra
|
||
governance -- "constrains" --> cognition
|
||
cognition -- "creates + refines" --> governance
|
||
```
|
||
|
||
### Entity relationship diagram — core entities and typed edges
|
||
|
||
```mermaid
|
||
erDiagram
|
||
MACHINE ||--o{ PROXMOX_HOST : "is-a"
|
||
PROXMOX_HOST ||--o{ LXC : hosts
|
||
PROXMOX_HOST ||--o{ VM : hosts
|
||
PROXMOX_HOST ||--o{ WORKSTATION : hosts
|
||
LXC ||--o{ SERVICE : provides
|
||
VM ||--o{ SERVICE : provides
|
||
WORKSTATION ||--o{ SERVICE : provides
|
||
LXC ||--o{ MOUNT : has
|
||
MOUNT }o--|| STORAGE_POOL : "stores-on"
|
||
SERVICE ||--o{ INGRESS_ROUTE : "exposed-by"
|
||
INGRESS_ROUTE }o--|| CERTIFICATE : "secured-by"
|
||
INGRESS_ROUTE }o--|| IDENTITY_PROVIDER : "secured-by"
|
||
SERVICE ||--o{ SERVICE : "depends-on"
|
||
CONFIG_REPO ||--|| SERVICE : "configured-by"
|
||
DEPLOY_PIPELINE ||--|| SERVICE : "deploys-to"
|
||
SERVICE ||--o{ SIGNAL : "monitored-by"
|
||
SIGNAL ||--o{ EXECUTION : triggers
|
||
EXECUTION ||--|| FEEDBACK : produces
|
||
FEEDBACK }o--|| PATTERN : "contributes-to"
|
||
PATTERN }o--|| SKILL : "informs"
|
||
SKILL ||--o{ CLASSIFICATION : "guides"
|
||
CLASSIFICATION ||--|| EXECUTION : "precedes"
|
||
EXECUTION ||--o| APPROVAL : "requires"
|
||
PERSON ||--o{ APPROVAL : "decides"
|
||
DOCUMENT ||--o{ SERVICE : "describes"
|
||
RUNBOOK ||--o{ SERVICE : "procedure-for"
|
||
AGENT ||--o{ EXECUTION : "performs"
|
||
PERSON ||--o{ AGENT : "owns"
|
||
```
|
||
|
||
### Infrastructure lifecycle
|
||
|
||
```mermaid
|
||
stateDiagram-v2
|
||
[*] --> planned : operator creates entity
|
||
planned --> provisioning : IP reserved, storage chosen, doc stub
|
||
provisioning --> active : mesh joined, health-check answering, doc complete
|
||
active --> migrating : preflight + backup verified
|
||
migrating --> active : post-verify, caddy checked, mounts checked
|
||
active --> deprecated : replacement live or role retired
|
||
deprecated --> destroyed : backups verified, secrets revoked, ingress removed
|
||
destroyed --> [*] : archaeology entry recorded
|
||
|
||
note right of deprecated
|
||
Complete only when
|
||
zero inbound edges remain
|
||
(no depends-on, no routes-to)
|
||
end note
|
||
```
|
||
|
||
### Signal lifecycle
|
||
|
||
```mermaid
|
||
stateDiagram-v2
|
||
[*] --> raised : scheduler probe or agent finding
|
||
raised --> acknowledged : agent or operator sees it
|
||
raised --> muted : operator suppresses (TTL)
|
||
acknowledged --> acting : actuator starts execution
|
||
acting --> resolved : action succeeded, verification passed
|
||
acting --> raised : action failed, re-escalated
|
||
raised --> resolved : condition cleared (auto-resolve)
|
||
muted --> raised : TTL expired
|
||
resolved --> [*]
|
||
```
|
||
|
||
### Execution + learning lifecycle
|
||
|
||
```mermaid
|
||
stateDiagram-v2
|
||
[*] --> proposed : signal + recommended_action
|
||
proposed --> approved : operator approves (if gated)
|
||
proposed --> auto_approved : risk class allows auto-act
|
||
approved --> executing : actuator runs
|
||
auto_approved --> executing : actuator runs
|
||
executing --> verified : verification command passed
|
||
executing --> failed : execution or verification failed
|
||
executing --> timed_out : exceeded duration limit
|
||
failed --> rolled_back : rollback procedure executed
|
||
verified --> [*] : feedback recorded, pattern updated
|
||
failed --> [*] : feedback recorded, pattern updated
|
||
rolled_back --> [*] : feedback recorded, pattern updated
|
||
timed_out --> [*] : feedback recorded, pattern updated
|
||
```
|
||
|
||
### Feedback / learning model — the cognition loop
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
subgraph observe["Observe"]
|
||
SIG["Signal raised\n(service down, disk full, drift)"]
|
||
end
|
||
|
||
subgraph decide["Decide"]
|
||
CLASS["Classification\nrisk × blast radius × confidence"]
|
||
SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"]
|
||
CLASS --> SKILL_LOOKUP
|
||
end
|
||
|
||
subgraph act["Act"]
|
||
EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"]
|
||
VERIFY["Verification\ncheck if action succeeded"]
|
||
EXEC --> VERIFY
|
||
end
|
||
|
||
subgraph learn["Learn (feedback loop)"]
|
||
OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"]
|
||
FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"]
|
||
PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"]
|
||
SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"]
|
||
|
||
OUTCOME --> FEEDBACK
|
||
FEEDBACK --> PATTERN
|
||
PATTERN --> SKILL_REFINE
|
||
end
|
||
|
||
SIG --> CLASS
|
||
SKILL_LOOKUP --> EXEC
|
||
VERIFY --> OUTCOME
|
||
SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP
|
||
|
||
subgraph escalation["Escalation path"]
|
||
APPROVAL["Approval request\n→ Matrix ✅/❌"]
|
||
end
|
||
|
||
CLASS -- "needs approval" --> APPROVAL
|
||
APPROVAL -- "approved" --> EXEC
|
||
APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"]
|
||
```
|
||
|
||
### Policy model — how risk attaches to entities
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
subgraph ontology["Ontology (DB metadata)"]
|
||
ET["entity_types\nhost, service, lxc, vm, ..."]
|
||
RT["relationship_types\nhosts, provides, depends-on, ..."]
|
||
LD["lifecycle_defs\nstates + transitions"]
|
||
end
|
||
|
||
subgraph policy["Policy (DB records)"]
|
||
RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"]
|
||
RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"]
|
||
AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"]
|
||
end
|
||
|
||
subgraph instances["Instances (DB data)"]
|
||
ENT["entities\nhost:hubris, service:caddy, ..."]
|
||
REL["relationships\nhubris hosts lxc:apps"]
|
||
end
|
||
|
||
ET --> RULES
|
||
RC --> RULES
|
||
RULES --> ENT
|
||
AUTO --> ENT
|
||
ET --> ENT
|
||
RT --> REL
|
||
LD --> ENT
|
||
```
|
||
|
||
## Target architecture
|
||
|
||
### Container stack on mac-mini
|
||
|
||
```mermaid
|
||
graph TB
|
||
subgraph mac-mini["mac-mini — Docker host, always-on"]
|
||
subgraph services["Docker Compose — Go binaries"]
|
||
PG["PostgreSQL 16\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata"]
|
||
INF["Infisical\n(secrets manager)"]
|
||
HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"]
|
||
API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events\n\nPolicy enforcement + risk\nclassification + ledger"]
|
||
SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• SSH to hubris/strong"]
|
||
NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"]
|
||
end
|
||
DEPLOY["Gitea webhook →\ndocker compose build + up -d"]
|
||
end
|
||
|
||
HERMES -- MCP --> API
|
||
API --> PG
|
||
SCHED --> PG
|
||
SCHED -- SSH --> HUBRIS
|
||
SCHED -- SSH --> STRONG
|
||
API -- "escalate" --> NOTIFIER
|
||
NOTIFIER -- "alerts + approvals" --> MATRIX
|
||
DEPLOY -- "rebuild" --> services
|
||
|
||
subgraph external["External"]
|
||
CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"]
|
||
APPS["apps/105 (fallback)"]
|
||
HUBRIS["hubris (PVE)"]
|
||
STRONG["strong (PVE)"]
|
||
GITEA["Gitea (LXC 104)"]
|
||
MATRIX["Matrix (LXC 118)"]
|
||
WS["Any workstation\nHermes remote → gateway"]
|
||
end
|
||
|
||
CADDY -- reverse_proxy --> API
|
||
GITEA -- webhook --> DEPLOY
|
||
WS -- "Hermes gateway" --> HERMES
|
||
```
|
||
|
||
### OODA loop — with learning feedback
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"]
|
||
ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"]
|
||
DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"]
|
||
DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"]
|
||
APPROVE -- "approved" --> ACT
|
||
ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"]
|
||
LEARN -. "improves confidence" .-> DECIDE
|
||
LEARN --> OBSERVE
|
||
```
|
||
|
||
### Deploy flow
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"]
|
||
GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"]
|
||
MACMINI -- "git pull" --> REPO["repo clone"]
|
||
MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"]
|
||
STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"]
|
||
```
|
||
|
||
## DB-native configuration
|
||
|
||
The three YAML files — `inventory.yaml`, `ontology.yaml`, `policy.yaml` — become
|
||
**seed manifests**. They bootstrap the DB on first deploy. After that, the DB is the
|
||
runtime source of truth, editable via the API. A future frontend can edit all three
|
||
directly.
|
||
|
||
### How it works
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
subgraph seeds["Seed manifests (git-tracked, YAML)"]
|
||
ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"]
|
||
INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"]
|
||
POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"]
|
||
end
|
||
|
||
subgraph db["PostgreSQL (runtime source of truth)"]
|
||
META["entity_types table\nrelationship_types table\nlifecycle_defs table"]
|
||
INST["entities table\nrelationships table"]
|
||
POLDB["policies table\nrisk_classes table\nautonomy_settings table"]
|
||
end
|
||
|
||
INGEST["Seed ingest (on deploy)\nidempotent upsert"]
|
||
ONTO_YAML --> INGEST --> META
|
||
INV_YAML --> INGEST --> INST
|
||
POL_YAML --> INGEST --> POLDB
|
||
|
||
API_EDIT["API edits\n(POST/PUT/PATCH)"]
|
||
API_EDIT --> META
|
||
API_EDIT --> INST
|
||
API_EDIT --> POLDB
|
||
|
||
EXPORT["Export to YAML\n(for DR / version control)"]
|
||
META --> EXPORT
|
||
INST --> EXPORT
|
||
POLDB --> EXPORT
|
||
```
|
||
|
||
### Why DB-native
|
||
|
||
- **Querying** — the agent can ask "what services depend on authentik?" as a graph
|
||
query, not a YAML parse. Blast-radius walks are SQL, not file reads.
|
||
- **Mutation** — adding a service, updating a lifecycle state, changing a policy rule
|
||
are DB transactions with audit trail, not file edits + git commits.
|
||
- **Consistency** — the ontology, inventory, and policy are always in sync (same DB,
|
||
same transaction). No drift between what the YAML says and what the runtime sees.
|
||
- **Future frontend** — a UI can edit entities, relationships, and policies directly
|
||
via the API. No need to generate/edit YAML files.
|
||
- **Version control** — seed YAML files are still git-tracked for bootstrap and DR.
|
||
The API can export the current DB state back to YAML for commit.
|
||
|
||
## Repo layout (Go project)
|
||
|
||
```
|
||
/ # repo root
|
||
├── docker-compose.yml # the OS stack definition
|
||
├── Makefile # build, test, deploy targets
|
||
├── go.mod # Go module definition
|
||
├── go.sum
|
||
├── cmd/ # binary entrypoints (one per service)
|
||
│ ├── api/ # Oikos API server
|
||
│ │ └── main.go
|
||
│ ├── scheduler/ # Observe + Act loop
|
||
│ │ └── main.go
|
||
│ └── notifier/ # Notification service
|
||
│ └── main.go
|
||
├── internal/ # private packages (not importable)
|
||
│ ├── db/ # database layer
|
||
│ │ ├── queries/ # sqlc SQL queries
|
||
│ │ ├── models.go # generated Go types
|
||
│ │ └── db.go # connection pool, migrations
|
||
│ ├── ontology/ # ontology types + meta-schema
|
||
│ │ ├── types.go # EntityType, RelationshipType, LifecycleDef
|
||
│ │ ├── graph.go # graph traversal (blast radius, dependencies)
|
||
│ │ └── ingest.go # YAML seed → DB ingest
|
||
│ ├── api/ # HTTP + MCP server
|
||
│ │ ├── server.go # Gin app setup
|
||
│ │ ├── routes/ # REST handlers
|
||
│ │ │ ├── hosts.go
|
||
│ │ │ ├── services.go
|
||
│ │ │ ├── signals.go
|
||
│ │ │ ├── approvals.go
|
||
│ │ │ ├── exec.go
|
||
│ │ │ └── knowledge.go
|
||
│ │ └── mcp.go # MCP protocol adapter (JSON-RPC over SSE)
|
||
│ ├── policy/ # risk classification + approval
|
||
│ │ ├── classify.go # risk × blast × confidence
|
||
│ │ ├── approve.go # approval request + grant lifecycle
|
||
│ │ └── autonomy.go # kill-switch, never-auto-act list
|
||
│ ├── scheduler/ # Observe stage
|
||
│ │ ├── probe.go # HTTP health, disk, drift
|
||
│ │ └── signal.go # raise/resolve signals in DB
|
||
│ ├── actuator/ # Act stage
|
||
│ │ ├── act.go # read signals, classify, execute or escalate
|
||
│ │ ├── execute.go # SSH execution + verification
|
||
│ │ └── guard.go # loop-guard, retry caps
|
||
│ ├── learning/ # feedback loop (the learning model)
|
||
│ │ ├── feedback.go # record outcome + lesson from execution
|
||
│ │ ├── pattern.go # extract/validate patterns from feedback
|
||
│ │ └── skill.go # create/refine skills from patterns
|
||
│ ├── notifier/ # notification abstraction
|
||
│ │ ├── notifier.go # interface
|
||
│ │ └── matrix.go # Matrix implementation
|
||
│ └── config/ # config loading (env, files)
|
||
│ └── config.go
|
||
├── migrations/ # SQL migrations (golang-migrate format)
|
||
│ ├── 001_ontology.up.sql # meta-schema (entity_types, relationship_types, ...)
|
||
│ ├── 001_ontology.down.sql
|
||
│ ├── 002_instances.up.sql # entities, relationships
|
||
│ ├── 003_operations.up.sql # signals, approvals, executions, feedback
|
||
│ ├── 004_learning.up.sql # patterns, skills
|
||
│ └── 005_policy.up.sql # policies, risk_classes, autonomy
|
||
├── seeds/ # YAML seed manifests (bootstrap + DR)
|
||
│ ├── ontology.yaml # entity types, relationship types, lifecycles
|
||
│ ├── inventory.yaml # entity instances (hosts, services, etc.)
|
||
│ └── policy.yaml # risk classes, approval rules, autonomy
|
||
├── docs/ # narrative docs (ingested into knowledge graph)
|
||
│ ├── containers/
|
||
│ ├── hosts/
|
||
│ ├── infrastructure/
|
||
│ └── investigations/
|
||
├── hermes/ # Hermes agent config + skills
|
||
│ ├── config.yaml
|
||
│ ├── SOUL.md
|
||
│ └── skills/
|
||
│ └── homelab-ops/
|
||
│ └── SKILL.md
|
||
├── compose/ # Docker build contexts
|
||
│ ├── api/Dockerfile
|
||
│ ├── scheduler/Dockerfile
|
||
│ ├── hermes/Dockerfile
|
||
│ └── postgres/init.sql
|
||
└── scripts/ # utility scripts
|
||
├── migrate-sops.sh # one-time SOPS → Infisical migration
|
||
└── import-legacy.sh # import existing signals/ledger JSONL
|
||
```
|
||
|
||
## Database schema
|
||
|
||
The schema is the ontology made concrete. Five migration groups, each adding a layer.
|
||
|
||
### Migration 1: Ontology meta-schema
|
||
|
||
```sql
|
||
-- The meta-graph: defines what entity types and relationship types can exist.
|
||
-- This IS the ontology, stored in the DB, editable via API.
|
||
|
||
CREATE TABLE entity_types (
|
||
name TEXT PRIMARY KEY, -- 'host', 'service', 'signal', 'pattern'
|
||
domain TEXT NOT NULL, -- 'physical', 'compute', 'network', ...
|
||
layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition'
|
||
description TEXT,
|
||
lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle)
|
||
created_at TIMESTAMPTZ DEFAULT now(),
|
||
updated_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
|
||
CREATE TABLE relationship_types (
|
||
name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on'
|
||
inverse TEXT, -- 'runs-on', 'provided-by'
|
||
source_type TEXT REFERENCES entity_types(name),
|
||
target_type TEXT REFERENCES entity_types(name),
|
||
cardinality TEXT NOT NULL, -- 'one-to-one', 'one-to-many', 'many-to-many'
|
||
description TEXT,
|
||
created_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
|
||
CREATE TABLE lifecycle_defs (
|
||
id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', 'execution', ...
|
||
states TEXT[] NOT NULL, -- ordered states
|
||
default_state TEXT NOT NULL,
|
||
transitions JSONB NOT NULL, -- {"from": {"to": {"requires": [...]}}}
|
||
created_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
```
|
||
|
||
### Migration 2: Entity instances (the inventory graph)
|
||
|
||
```sql
|
||
-- Entity instances — the actual hosts, services, signals, patterns, etc.
|
||
-- This replaces inventory.yaml as the runtime source of truth.
|
||
|
||
CREATE TABLE entities (
|
||
id TEXT PRIMARY KEY, -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001'
|
||
type TEXT NOT NULL REFERENCES entity_types(name),
|
||
name TEXT NOT NULL, -- 'hubris', 'caddy', 'disk-threshold'
|
||
state TEXT, -- lifecycle state (e.g. 'active', 'raised')
|
||
attributes JSONB NOT NULL DEFAULT '{}', -- type-specific data (IP, mesh addr, port, ...)
|
||
parent_id TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host)
|
||
created_at TIMESTAMPTZ DEFAULT now(),
|
||
updated_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
|
||
CREATE INDEX idx_entities_type ON entities(type);
|
||
CREATE INDEX idx_entities_state ON entities(state);
|
||
CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes);
|
||
|
||
-- Relationship instances — the typed edges of the graph
|
||
CREATE TABLE relationships (
|
||
source_id TEXT NOT NULL REFERENCES entities(id),
|
||
target_id TEXT NOT NULL REFERENCES entities(id),
|
||
type TEXT NOT NULL REFERENCES relationship_types(name),
|
||
attributes JSONB,
|
||
created_at TIMESTAMPTZ DEFAULT now(),
|
||
PRIMARY KEY (source_id, target_id, type)
|
||
);
|
||
|
||
CREATE INDEX idx_rel_source ON relationships(source_id);
|
||
CREATE INDEX idx_rel_target ON relationships(target_id);
|
||
CREATE INDEX idx_rel_type ON relationships(type);
|
||
|
||
-- Recursive graph traversal function (blast radius, dependency chains)
|
||
CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3)
|
||
RETURNS TABLE(entity_id TEXT, depth INT) AS $$
|
||
WITH RECURSIVE walk AS (
|
||
SELECT start_id::TEXT AS entity_id, 0::INT AS depth
|
||
UNION
|
||
SELECT r.target_id::TEXT, w.depth + 1
|
||
FROM relationships r
|
||
JOIN walk w ON r.source_id = w.entity_id
|
||
WHERE w.depth < max_depth
|
||
)
|
||
SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
|
||
$$ LANGUAGE sql STABLE;
|
||
```
|
||
|
||
### Migration 3: Operations (signals, approvals, ledger, state)
|
||
|
||
```sql
|
||
-- Signals — now entities in the graph, with a dedicated table for indexed querying
|
||
-- (the entity row is the canonical record; this table is a fast lookup)
|
||
CREATE TABLE signals (
|
||
entity_id TEXT PRIMARY KEY REFERENCES entities(id),
|
||
kind TEXT NOT NULL,
|
||
severity TEXT NOT NULL, -- info, warning, critical
|
||
target_entity_id TEXT REFERENCES entities(id), -- the infrastructure entity this is about
|
||
evidence TEXT,
|
||
likely_cause TEXT,
|
||
recommended_action JSONB,
|
||
verification TEXT,
|
||
state TEXT NOT NULL DEFAULT 'raised',
|
||
mute_until TIMESTAMPTZ,
|
||
created_at TIMESTAMPTZ DEFAULT now(),
|
||
updated_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
|
||
CREATE INDEX idx_signals_state ON signals(state);
|
||
CREATE INDEX idx_signals_target ON signals(target_entity_id);
|
||
CREATE INDEX idx_signals_severity ON signals(severity);
|
||
|
||
-- Approvals
|
||
CREATE TABLE approvals (
|
||
id TEXT PRIMARY KEY,
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
entity_id TEXT REFERENCES entities(id), -- entity to act on
|
||
action TEXT NOT NULL,
|
||
risk_class TEXT NOT NULL,
|
||
status TEXT NOT NULL DEFAULT 'pending', -- pending, approved, denied, expired
|
||
ttl INTERVAL NOT NULL DEFAULT '1 hour',
|
||
decided_at TIMESTAMPTZ,
|
||
decided_by TEXT REFERENCES entities(id), -- person entity
|
||
confirmation_phrase TEXT
|
||
);
|
||
|
||
-- Change ledger (high-level record, links to execution for detail)
|
||
CREATE TABLE ledger_entries (
|
||
id SERIAL PRIMARY KEY,
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
entity_id TEXT REFERENCES entities(id),
|
||
action TEXT NOT NULL,
|
||
risk_class TEXT NOT NULL,
|
||
result TEXT, -- ok, failed, escalated
|
||
approval_id TEXT REFERENCES approvals(id),
|
||
execution_id INTEGER, -- FK to executions (migration 4)
|
||
agent_id TEXT REFERENCES entities(id),
|
||
notes TEXT
|
||
);
|
||
|
||
-- State snapshots (replaces oikos/state.json)
|
||
CREATE TABLE state_snapshots (
|
||
id SERIAL PRIMARY KEY,
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
entity_id TEXT REFERENCES entities(id),
|
||
health TEXT, -- healthy, degraded, down, unknown
|
||
data JSONB
|
||
);
|
||
```
|
||
|
||
### Migration 4: Learning model (executions, feedback, patterns, skills)
|
||
|
||
```sql
|
||
-- Executions — detailed record of each action the OS performs
|
||
CREATE TABLE executions (
|
||
id SERIAL PRIMARY KEY,
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
signal_entity_id TEXT REFERENCES entities(id), -- signal that triggered this
|
||
target_entity_id TEXT REFERENCES entities(id), -- entity acted upon
|
||
action TEXT NOT NULL,
|
||
risk_class TEXT NOT NULL,
|
||
approval_id TEXT REFERENCES approvals(id),
|
||
agent_id TEXT REFERENCES entities(id), -- who/what executed
|
||
skill_id TEXT REFERENCES entities(id), -- skill used (if any)
|
||
status TEXT NOT NULL DEFAULT 'queued', -- queued, running, completed, failed, timed-out
|
||
result JSONB, -- detailed result data
|
||
duration_ms INTEGER,
|
||
verified BOOLEAN DEFAULT false,
|
||
started_at TIMESTAMPTZ,
|
||
completed_at TIMESTAMPTZ
|
||
);
|
||
|
||
CREATE INDEX idx_exec_target ON executions(target_entity_id);
|
||
CREATE INDEX idx_exec_status ON executions(status);
|
||
CREATE INDEX idx_exec_action ON executions(action);
|
||
|
||
-- Feedback — what was learned from an execution
|
||
CREATE TABLE feedback (
|
||
id SERIAL PRIMARY KEY,
|
||
execution_id INTEGER REFERENCES executions(id),
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
outcome TEXT NOT NULL, -- success, failure, partial, unexpected
|
||
observation TEXT, -- what happened vs what was expected
|
||
lesson TEXT, -- extractable lesson
|
||
unexpected_side_effects TEXT[],
|
||
tags TEXT[]
|
||
);
|
||
|
||
CREATE INDEX idx_feedback_execution ON feedback(execution_id);
|
||
CREATE INDEX idx_feedback_outcome ON feedback(outcome);
|
||
|
||
-- Patterns — generalized rules extracted from accumulated feedback
|
||
CREATE TABLE patterns (
|
||
id TEXT PRIMARY KEY, -- 'pat-2026-07-06-001'
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
entity_type TEXT REFERENCES entity_types(name), -- applies to this type
|
||
action TEXT NOT NULL, -- 'restart', 'deploy', etc.
|
||
pattern TEXT NOT NULL, -- 'service X recovers within 30s after restart'
|
||
confidence REAL DEFAULT 0.5, -- 0.0 to 1.0
|
||
evidence_count INTEGER DEFAULT 1, -- how many executions support this
|
||
success_count INTEGER DEFAULT 0,
|
||
failure_count INTEGER DEFAULT 0,
|
||
status TEXT DEFAULT 'hypothesized', -- hypothesized, validated, active, deprecated
|
||
last_validated_at TIMESTAMPTZ
|
||
);
|
||
|
||
CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action);
|
||
CREATE INDEX idx_patterns_status ON patterns(status);
|
||
|
||
-- Skills — codified procedures refined through feedback
|
||
CREATE TABLE skills (
|
||
id TEXT PRIMARY KEY, -- 'skill-restart-service', 'skill-deploy-lxc'
|
||
name TEXT NOT NULL,
|
||
ts TIMESTAMPTZ DEFAULT now(),
|
||
procedure TEXT NOT NULL, -- the codified steps (markdown or structured)
|
||
applies_to TEXT REFERENCES entity_types(name),
|
||
pattern_ids TEXT[], -- patterns that inform this skill
|
||
status TEXT DEFAULT 'drafted', -- drafted, tested, active, refined, deprecated
|
||
version INTEGER DEFAULT 1,
|
||
success_rate REAL, -- rolling success rate
|
||
last_used_at TIMESTAMPTZ
|
||
);
|
||
|
||
CREATE INDEX idx_skills_type ON skills(applies_to);
|
||
CREATE INDEX idx_skills_status ON skills(status);
|
||
```
|
||
|
||
### Migration 5: Policy (DB-native risk + approval rules)
|
||
|
||
```sql
|
||
-- Risk classes — the four-level safety model
|
||
CREATE TABLE risk_classes (
|
||
name TEXT PRIMARY KEY, -- 'read_only', 'reversible_low', etc.
|
||
description TEXT,
|
||
approval_required TEXT NOT NULL DEFAULT 'none', -- none, operator, operator_confirmed
|
||
ledger BOOLEAN DEFAULT false,
|
||
autonomy_allowed BOOLEAN DEFAULT false -- can agent auto-act at this risk level?
|
||
);
|
||
|
||
-- Approval rules — entity_type + action → risk_class + requirements
|
||
CREATE TABLE approval_rules (
|
||
id SERIAL PRIMARY KEY,
|
||
entity_type TEXT REFERENCES entity_types(name), -- applies to this entity type
|
||
action TEXT NOT NULL, -- 'restart', 'deploy', 'destroy'
|
||
risk_class TEXT NOT NULL REFERENCES risk_classes(name),
|
||
autonomy_level TEXT NOT NULL DEFAULT 'auto', -- 'auto', 'escalate', 'never'
|
||
scope_entity TEXT REFERENCES entities(id), -- optional: specific entity only
|
||
created_at TIMESTAMPTZ DEFAULT now(),
|
||
updated_at TIMESTAMPTZ DEFAULT now(),
|
||
UNIQUE(entity_type, action)
|
||
);
|
||
|
||
-- Autonomy settings — global kill-switch + per-entity overrides
|
||
CREATE TABLE autonomy_settings (
|
||
key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.caddy'
|
||
value TEXT NOT NULL, -- 'off', 'reversible_low', 'true', 'false'
|
||
updated_at TIMESTAMPTZ DEFAULT now()
|
||
);
|
||
```
|
||
|
||
## Workstreams
|
||
|
||
### 1. Ontology definition + seed manifests (`seeds/`, `internal/ontology/`)
|
||
|
||
**Before any code**, finalize the ontology. The current `ontology.yaml` has 8 domains
|
||
and 14 relationship types. The new ontology adds:
|
||
|
||
- **Layer 3 entities** (cognition): signal, execution, feedback, pattern, skill,
|
||
approval, classification, document, runbook
|
||
- **New relationships**: `triggers`, `produces`, `contributes-to`, `informs`, `guides`,
|
||
`precedes`, `performs`, `procedure-for`, `learned-from`
|
||
- **Lifecycle definitions** for operational entities (signals, executions, patterns,
|
||
skills) — not just infrastructure
|
||
|
||
**Deliverables:**
|
||
- `seeds/ontology.yaml` — entity types, relationship types, lifecycle definitions
|
||
(seeded into `entity_types`, `relationship_types`, `lifecycle_defs` on deploy)
|
||
- `seeds/inventory.yaml` — adapted from current `inventory.yaml` (seeded into
|
||
`entities` + `relationships`)
|
||
- `seeds/policy.yaml` — adapted from current `policy.yaml` (seeded into `risk_classes`,
|
||
`approval_rules`, `autonomy_settings`)
|
||
- `internal/ontology/ingest.go` — idempotent seed → DB ingest
|
||
|
||
### 2. Database layer + migrations (`migrations/`, `internal/db/`)
|
||
|
||
- Write the 5 migrations above
|
||
- Set up sqlc for type-safe Go database access
|
||
- Connection pool, migration runner
|
||
- Graph traversal queries (blast radius, dependency chains, knowledge lookup)
|
||
- Import script for existing `signals/*.jsonl`, `ledger/*.jsonl` → DB
|
||
|
||
### 3. Unified API server — Go (`cmd/api/`, `internal/api/`)
|
||
|
||
One Go binary (Gin web framework) exposing REST + MCP from the same codebase.
|
||
|
||
**MCP interface** (`internal/api/mcp.go`):
|
||
- JSON-RPC over SSE, compatible with Hermes MCP client
|
||
- Tools: `get_host`, `list_services`, `search_knowledge`, `get_entity`,
|
||
`get_relations`, `get_blast_radius`, `get_signal_history`, `get_ledger`,
|
||
`get_patterns`, `get_skills`, `get_state_snapshot`
|
||
- All read from PostgreSQL
|
||
|
||
**REST interface** (`internal/api/routes/`):
|
||
- `GET /api/v1/entities` — list entities (filter by type, state, domain)
|
||
- `GET /api/v1/entities/{id}` — entity detail + relationships
|
||
- `POST /api/v1/entities` — create entity (creates inventory entry)
|
||
- `PATCH /api/v1/entities/{id}` — update entity (state transition, attributes)
|
||
- `GET /api/v1/signals` — list signals (filter by state, severity, entity)
|
||
- `POST /api/v1/signals/{id}/ack` — acknowledge
|
||
- `POST /api/v1/signals/{id}/resolve` — resolve
|
||
- `GET /api/v1/approvals` — pending approvals
|
||
- `POST /api/v1/approvals/{id}/decide` — approve/deny (Authentik-gated)
|
||
- `POST /api/v1/exec` — gated execution (classify → check approval → execute → feedback)
|
||
- `GET /api/v1/patterns` — list patterns (filter by entity_type, action, status)
|
||
- `GET /api/v1/skills` — list skills (filter by applies_to, status)
|
||
- `GET /api/v1/knowledge/{entity_id}` — knowledge graph query
|
||
- `GET /api/v1/knowledge/search?q=...` — search knowledge graph
|
||
- `GET /api/v1/ontology` — list entity types, relationship types, lifecycles
|
||
- `POST /api/v1/ontology/entity_types` — create entity type (extend the schema)
|
||
- `WS /api/v1/events` — real-time stream (signals, approvals, executions, feedback)
|
||
|
||
**Policy enforcement** (`internal/policy/classify.go`):
|
||
- Every mutating endpoint classifies the action via the policy DB
|
||
- Risk class → approval check → autonomy check
|
||
- All mutations write to the ledger automatically
|
||
|
||
**Auth:**
|
||
- MCP interface: no auth (internal, container-to-container)
|
||
- REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints
|
||
- Internal: shared secret (Docker network)
|
||
|
||
### 4. Scheduler + Actuator — Go (`cmd/scheduler/`, `internal/scheduler/`, `internal/actuator/`)
|
||
|
||
**Scheduler (Observe)** — Go service with goroutines for concurrent probes:
|
||
- HTTP health probes (concurrent, with timeouts)
|
||
- Disk usage probes (SSH to hubris/strong)
|
||
- Drift detection (inventory vs live state)
|
||
- Writes signals + state snapshots to DB
|
||
- Runs on a 10-min ticker
|
||
|
||
**Actuator (Act)** — Go service, the control loop:
|
||
- Reads open signals with `recommended_action`
|
||
- For each: classify via `internal/policy/classify.go`
|
||
- **auto-act**: look up skill for (entity_type, action) → follow procedure → execute
|
||
via SSH → verify → record execution → generate feedback → update patterns
|
||
- **escalate**: create approval request → notify via Matrix → acknowledge signal
|
||
- Loop-guard: check execution history per (entity, action) to cap auto-retries
|
||
(`SELECT ... FOR UPDATE SKIP LOCKED` for concurrency safety)
|
||
- Autonomy kill-switch: check `autonomy_settings` table
|
||
|
||
### 5. Learning engine — Go (`internal/learning/`)
|
||
|
||
The feedback loop that makes the agent improve over time.
|
||
|
||
**Feedback recording** (`internal/learning/feedback.go`):
|
||
- After every execution, evaluate the outcome:
|
||
- Did the verification command pass? → success
|
||
- Did it fail? → failure
|
||
- Did it partially work? → partial
|
||
- Did something unexpected happen? → unexpected
|
||
- Record a feedback entry with: outcome, observation (what happened vs expected),
|
||
lesson (extractable insight), unexpected_side_effects
|
||
|
||
**Pattern extraction** (`internal/learning/pattern.go`):
|
||
- Periodically scan accumulated feedback for (entity_type, action) pairs
|
||
- When N+ executions share a similar outcome, extract a pattern:
|
||
- "Restarting service:X typically takes 15s and succeeds"
|
||
- "Deploying to LXC:Y via webhook has 30% failure rate, retry helps"
|
||
- Patterns start as `hypothesized`, move to `validated` after enough evidence,
|
||
then `active` (used by the decision classifier)
|
||
- Confidence score = success_count / evidence_count, adjusted by recency
|
||
|
||
**Skill management** (`internal/learning/skill.go`):
|
||
- When a pattern reaches `active` status with confidence > 0.7, create or refine
|
||
a skill for that (entity_type, action) pair
|
||
- Skills codify the best-known procedure (what steps to take, what to verify,
|
||
expected duration, known failure modes)
|
||
- Skills are versioned — each refinement increments the version
|
||
- The actuator looks up skills before executing: if a skill exists, follow it;
|
||
if not, use the default procedure and generate feedback for future pattern extraction
|
||
|
||
**How the classifier uses learning** (`internal/policy/classify.go`):
|
||
- Confidence scoring now checks patterns + skills, not just raw ledger history:
|
||
- If a pattern exists for (entity_type, action) with high confidence → boost
|
||
auto-act confidence
|
||
- If patterns show frequent failures → lower confidence, escalate
|
||
- If a skill exists → higher confidence (proven procedure available)
|
||
- This is the closed loop: **execution → feedback → pattern → skill → classification
|
||
→ execution** (better informed each time)
|
||
|
||
### 6. Knowledge graph ingestion (`internal/ontology/ingest.go`)
|
||
|
||
After the ontology is defined and the DB schema is in place:
|
||
|
||
- On deploy, walk `docs/` directory
|
||
- Parse each markdown file:
|
||
- Extract frontmatter for metadata (entity type, tags, relations)
|
||
- Infer entity relationships from path conventions:
|
||
`docs/containers/105-apps.md` → relationship to `entity:lxc:apps`
|
||
- Extract cross-references (markdown links) → relationships
|
||
- Create knowledge entities in the `entities` table (type = `document`, `runbook`,
|
||
`investigation`, etc.) with `documents` / `procedure-for` edges to infrastructure
|
||
entities
|
||
- Idempotent — safe to re-run on every deploy
|
||
|
||
**Agent access:**
|
||
- MCP tool `search_knowledge(query)` — full-text search on knowledge entities
|
||
- MCP tool `get_entity_knowledge(entity_id)` — all docs related to an entity
|
||
- MCP tool `get_relations(entity_id)` — graph traversal (blast radius, dependencies)
|
||
|
||
### 7. Hermes agent container (`compose/hermes/`)
|
||
|
||
- Hermes Agent runtime in a Docker container, gateway mode
|
||
- Config: `hermes/config.yaml` (providers, models, gateway port)
|
||
- Persona: `hermes/SOUL.md` (homelab-specific)
|
||
- Skills: `hermes/skills/homelab-ops/SKILL.md` — how to use the Oikos API, classify
|
||
actions, request approvals, query the knowledge graph
|
||
- Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid)
|
||
- Gateway port 8092 — workstations connect remotely
|
||
- Hermes data volume for persistent state
|
||
|
||
### 8. Infisical secrets migration
|
||
|
||
Same as rev 1:
|
||
- Stand up Infisical in the Docker stack
|
||
- Migrate SOPS secrets (one-time decrypt + import)
|
||
- Wire all Go services to Infisical via machine identity
|
||
- Retire SOPS + age keys
|
||
|
||
### 9. Notifier — Go (`cmd/notifier/`, `internal/notifier/`)
|
||
|
||
**Interface (Go):**
|
||
```go
|
||
type Notifier interface {
|
||
SendAlert(ctx context.Context, signal Signal) error
|
||
SendApprovalRequest(ctx context.Context, approval Approval) error
|
||
ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error)
|
||
}
|
||
```
|
||
|
||
**Matrix implementation:**
|
||
- Sends alerts to `@dtoro:avispero` via Synapse (LXC 118)
|
||
- Approval requests as messages with ✅/❌ reactions
|
||
- Listens for reactions to record decisions
|
||
- Pluggable — future implementations (webhook, email) register via config
|
||
|
||
### 10. Docker build + deploy pipeline
|
||
|
||
- Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch)
|
||
- `docker compose build` from repo root
|
||
- Gitea webhook on push to `main` → deploy script on mac-mini
|
||
- Deploy script: `git pull && go build ./... && docker compose build && docker compose up -d`
|
||
- Seed ingest runs as part of the API startup (idempotent)
|
||
- Health checks on each service
|
||
|
||
### 11. Ingress re-point (Caddy)
|
||
|
||
- `dtoro/caddy-conf`: point `mcp.hubris.network` + `oikos.hubris.network` →
|
||
mac-mini mesh IP :8090 (API)
|
||
- Future: `hermes.hubris.network` → mac-mini:8092
|
||
- DNS and public URLs unchanged
|
||
|
||
### 12. mac-mini host setup
|
||
|
||
- New directory: `~/oikos-os/` — repo clone + `docker compose` working dir
|
||
- Existing `/opt/homelab-context/` stays untouched until cleanup phase
|
||
- Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys
|
||
- Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone
|
||
|
||
### 13. Decommission apps/105 (deferred)
|
||
|
||
Keep apps/105 running as fallback. Cutover checklist when ready:
|
||
1. Verify Docker OS serves all traffic
|
||
2. `systemctl disable --now` Oikos services on apps/105
|
||
3. Remove old checkouts
|
||
4. Update Caddy backends exclusively to mac-mini
|
||
5. Remove/retarget Gitea webhooks
|
||
|
||
## Phasing
|
||
|
||
**Phase 0 — Ontology design (no code):**
|
||
- Finalize entity types, relationship types, lifecycles
|
||
- Write seed manifests (`seeds/ontology.yaml`, `seeds/inventory.yaml`, `seeds/policy.yaml`)
|
||
- Review diagrams with operator
|
||
|
||
**Phase 1 — Foundation (Go + DB):**
|
||
- Go module setup, project structure
|
||
- PostgreSQL migrations (all 5)
|
||
- Seed ingest pipeline (YAML → DB)
|
||
- sqlc queries for core operations
|
||
- Import existing signals/ledger data
|
||
|
||
**Phase 2 — API (Go):**
|
||
- Gin server with REST routes
|
||
- MCP protocol adapter
|
||
- Policy enforcement middleware
|
||
- Knowledge graph ingestion from docs/
|
||
|
||
**Phase 3 — Control loop (Go):**
|
||
- Scheduler (Observe) — probes, signals, state snapshots
|
||
- Actuator (Act) — classify, execute, verify
|
||
- Learning engine — feedback, patterns, skills
|
||
|
||
**Phase 4 — Agent (Hermes container):**
|
||
- Hermes Docker image, gateway config
|
||
- Homelab skills
|
||
- Connect from workstation, verify MCP + SSH
|
||
|
||
**Phase 5 — Secrets (Infisical):**
|
||
- Stand up Infisical, migrate SOPS, wire services
|
||
|
||
**Phase 6 — Deploy + cutover:**
|
||
- Docker Compose, Gitea webhook, Caddy re-point
|
||
- End-to-end verification
|
||
- Stop apps/105, clean up mac-mini
|
||
|
||
## Reuse (logic carried over, rewritten in Go)
|
||
|
||
| Existing Python | What it becomes in Go |
|
||
|---|---|
|
||
| `oikos/decide.py` | `internal/policy/classify.go` — same scoring logic, reads from DB + patterns |
|
||
| `oikos/signal.py` | `internal/scheduler/signal.go` — same lifecycle, DB-backed |
|
||
| `oikos/approve.py` | `internal/policy/approve.go` — same grant lifecycle, DB-backed |
|
||
| `oikos/ledger.py` | `internal/db/` — ledger_entries table + sqlc queries |
|
||
| `oikos/policy.py` + `policy.yaml` | `internal/policy/` + `seeds/policy.yaml` → DB tables |
|
||
| `oikos/drift.py` | `internal/scheduler/` — drift detection, writes signals to DB |
|
||
| `oikos/relations.py` | `internal/ontology/graph.go` — SQL graph traversal |
|
||
| `oikos/report.py` | `internal/api/routes/` — report endpoints, reads from DB |
|
||
| `mcp/server.py` | `internal/api/mcp.go` — MCP adapter on top of DB |
|
||
| `bin/homelab` logic | `internal/api/routes/` — same operations, REST interface |
|
||
| `oikos/scheduler.py` | `internal/scheduler/` — same probes, goroutines for concurrency |
|
||
| `oikos/approve.py` Matrix delivery | `internal/notifier/matrix.go` |
|
||
| *(new)* | `internal/learning/` — feedback, patterns, skills (no Python equivalent) |
|
||
|
||
## Risks / trade-offs
|
||
|
||
- **Go rewrite** — the existing Python code (~4400 lines) is replaced. The logic and
|
||
design patterns carry over, but it's a full rewrite. Mitigated by the fact that the
|
||
Python code is well-documented and the Go structure mirrors it.
|
||
- **Hermes in Docker** — agent's world is the container. SSH access is the bridge.
|
||
Hybrid approach (mounted keys now, actuator gateway later).
|
||
- **PostgreSQL as SPOF** — mitigated by Docker volume persistence + automated
|
||
`pg_dump` backups (the scheduler can do this once running).
|
||
- **Learning model cold start** — no patterns/skills exist initially. The agent starts
|
||
cautious (escalates everything), accumulates feedback, and gradually becomes more
|
||
autonomous as patterns validate. This is by design — trust is earned.
|
||
- **Infisical bootstrapping** — SOPS coexists during transition. Keep SOPS as fallback.
|
||
- **Multi-agent concurrency** — `SELECT ... FOR UPDATE SKIP LOCKED` prevents two
|
||
actuator passes from acting on the same signal.
|
||
- **Ontology evolution** — as the homelab changes, entity types and relationship types
|
||
need to be added/modified. The DB-native approach makes this an API call, not a
|
||
file edit + redeploy.
|
||
|
||
## Verification (end to end)
|
||
|
||
1. **Ontology:** `seeds/ontology.yaml` ingested — `SELECT * FROM entity_types` shows
|
||
all types across 3 layers; lifecycle definitions match the state machine diagrams.
|
||
2. **DB:** `docker compose up postgres` — all 5 migrations applied; seed ingest
|
||
populates entities + relationships from `inventory.yaml`.
|
||
3. **API:** `curl http://localhost:8090/api/v1/entities?type=service` returns the
|
||
fleet; MCP `list_services` works via the same endpoint.
|
||
4. **Scheduler:** trigger a probe pass — signals in DB, state snapshots written.
|
||
5. **Actuator:** raise a test `service-down` signal → actuator classifies →
|
||
auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated.
|
||
6. **Learning:** after N executions of the same (entity_type, action), a pattern
|
||
appears with confidence score; after enough evidence, a skill is created.
|
||
7. **Classifier with learning:** set `autonomy.auto_act: reversible_low` → next
|
||
similar signal: classifier checks pattern confidence → auto-acts if high,
|
||
escalates if low. Kill-switch (`auto_act: off`) → always escalates.
|
||
8. **Hermes:** connect from another workstation → agent responds, queries API via
|
||
MCP, can SSH to hubris.
|
||
9. **Secrets:** Infisical running, Go services fetch secrets, SOPS files removed.
|
||
10. **Deploy:** `git push` → Gitea webhook → `docker compose build + up -d` →
|
||
changes live, seed ingest syncs any YAML changes to DB.
|
||
11. **Knowledge:** MCP `search_knowledge("caddy")` returns docs linked to
|
||
`entity:service:caddy`.
|
||
12. **Cutover:** stop apps/105, verify production traffic only from Docker OS.
|
||
|
||
## Out of scope (for now)
|
||
|
||
- Oikos Console web UI (deferred — built later on top of the API)
|
||
- Multi-node deployment (designed for, not implemented)
|
||
- Vector embeddings / semantic search (structured graph only for now)
|
||
- SSH-key-signed approval requests
|
||
- Prometheus / trend signals
|
||
- Actuator gateway pattern (Phase 2 of hybrid — start with mounted SSH)
|
||
- Automated skill extraction via LLM (patterns extracted statistically for now;
|
||
LLM-assisted skill refinement is a future enhancement)
|