Addresses 15 original HIGH audit findings + 18 new findings from systems architect + senior Go developer review (572 lines added). CRITICAL fixes: - SA1: Cognition objects (execution/feedback/pattern/skill) get dual entity pattern — entities row + typed table, graph-traversable - SG1: Hypertable PKs fixed — PRIMARY KEY (id, ts) for audit_log, events, agent_activity (was id-only, would fail create_hypertable) HIGH fixes: - SA2: Remove 'cognition creates governance' arrow (unsupported, was learning-poisoning vector). Patterns propose, operator accepts. - SA3: Add Person, Agent, IdentityProvider to ontology (were used in BDD but never defined) - SA4: Fix all lifecycle dead-ends — add 'failed' state to infra, terminal 'failed'/'invalidated' to signals/patterns/skills, add approval lifecycle diagram, add cancellation/rollback-failure to executions - SA5: Add classifications table — persist classifier reasoning (was modeled in BDD but never stored) - SA6: Move recommended_action from signals to classifications - SA8: Add Cluster, ComposeStack, ManagedHost to ontology - SG2: Drop array_agg from CAGG (unsupported by TimescaleDB) - SG3: Idempotent TimescaleDB calls (if_not_exists, exception guards) - SG4: Graceful shutdown (SIGTERM, in-flight protection, 30s grace) - SG5: Entity-level advisory locks (pg_advisory_xact_lock per target) - SG6: Domain layer (internal/domain/) — sqlc models never escape db/ Security: - S1: Restricted SSH key (command=) now + actuator gateway in Phase 3 - S2: MCP shared-secret auth + dedicated Docker network - S3: Policy mutations require meta-approval (dual-control) - S4: Pattern activation needs operator confirmation + confidence capped by sample size (N>=5) + anomaly detection - S5: Single-use HMAC approval tokens replace confirmation_phrase - SA10: Gateway mTLS + Caddy as documented trust root + JWT validation Operational: - A3/O3/O4: Backup to Proton Drive (daily pg_dump + WAL), restore runbook, DR plan (RTO 4h, RPO 24h), monthly restore drill - O1: Forward-only migrations + pre-deploy backup + rollback runbook - O2: External watchdog cron on apps/105 - M1: CI/CD via Gitea Actions (go vet, lint, test -race, docker build) Architecture: - A1: Testing strategy with specific tests per package + coverage gates - SA7: Notifier decoupled via DB rendezvous (no service-to-service calls) - SA9: TimescaleDB Docker image specified + init container for migrations - SG7: Pattern/skill management endpoints (operator override) - SG8: WebSocket push via in-process bus + LISTEN/NOTIFY - SG10: Transactional event emission (same tx as state change) - SG11: Error handling — sentinel errors + HTTP mapping + SSH taxonomy - SG13: Context-aware SSH (x/crypto/ssh doesn't honor context) - SG14: Connection pool sizing (28 total, max_connections=80) - SG15: RESTful /executions (was /exec) - SG16: Pagination on all list endpoints - SG17: Go tooling (sqlc.yaml, module path, CGO_ENABLED=0, distroless) - SG18: /healthz and /metrics bypass auth + audit Updated phasing incorporates all remediation.
115 KiB
Plan: Oikos — Docker-based agentic homelab OS on mac-mini
Status: Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first design with systems modeling, DB-native config (inventory/ontology/policy as graph metadata, not YAML files), and a feedback loop where the agent learns from execution.
Vision
Convert this repo into a Docker-based agentic homelab OS written in Go. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors:
- Operator (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix.
- Agent — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and learns from every action to improve over time.
All OS services run in Docker containers on mac-mini. The OS is deployed by a git push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path to multi-node later.
Decisions (from operator Q&A, 2026-07-06)
| Question | Decision |
|---|---|
| Repo structure | One repo, reorganize internally. |
| Language | Go — compiled, type-safe, small containers, goroutines for concurrent probes. |
| Hermes runtime | Runs inside Docker as part of the OS stack (gateway mode). |
| Agent → homelab access | Hybrid — mounted SSH keys now, actuator gateway built incrementally. |
| Data storage | PostgreSQL. Inventory/ontology/policy become DB-native graph metadata. |
| Config (inventory/ontology/policy) | DB-native — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend. |
| Knowledge/context | Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data. |
| Operator interface | Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks. |
| Host | mac-mini for now, designed to scale later. |
| Deploy | Git push → Gitea webhook → Docker rebuild + restart. |
bin/homelab CLI |
Replaced by an API. CLI becomes a thin Go client that calls the OS API. |
| Secrets | Migrate from SOPS+age to Infisical. |
| Matrix | Keep for now, abstract the notification layer for future channels. |
| MCP server | Merged into the unified API — one Go service, REST + MCP interfaces. |
| Agent type | Hermes core + custom homelab skills. |
| Feedback loop | Agent learns from execution — outcomes feed back as patterns and skills that inform future decisions. |
| Ontology | Developed first, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles. |
| apps/105 | Keep running as fallback until Docker OS is proven. |
| mac-mini cleanup | Start fresh in a new directory, clean up old artifacts later. |
Ontology — the systems model
The ontology is the foundational layer of the OS. It defines what exists, how things connect, how they change over time, and how the OS learns about them. It is stored in the database as metadata (entity types, relationship types, lifecycle definitions). YAML seed files bootstrap it on first deploy; after that, the DB is authoritative and editable via API.
Design principles (systems modeling)
- Three layers — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's own behavior + learning). Dependencies flow upward: Cognition depends on Governance depends on Infrastructure.
- Everything is an entity — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, changes, skills) are first-class entities, not second-class records.
- Typed relationships with cardinality — edges carry semantics.
hostsis one-to-many;depends-onis many-to-many;documentsis one-to-one. The graph is queryable for blast radius, dependency chains, and knowledge lookup. - Lifecycles are state machines — every entity type has a lifecycle. Infrastructure
entities move through
planned → active → destroyed. Operational entities have their own lifecycles (signals, approvals, patterns, skills). Transitions can require preconditions. - Policy is attached to the ontology — risk classes and approval rules link to entity types and actions. The policy IS part of the model, not a separate file.
- Learning is modeled — executions produce outcomes, outcomes accumulate into patterns, patterns refine skills, skills inform future decisions. This is an explicit, queryable part of the graph.
Layer map
graph TB
cognition -- "observes, acts on, learns about" --> infra
governance -- "governs access to" --> infra
governance -- "constrains" --> cognition
cognition -- "creates + refines" --> governance
subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"]
direction LR
OBS["Observation\nsignal, state-snapshot"]
DEC["Decision\nclassification, risk-assessment"]
ACT["Action\nexecution, verification"]
GOV["Governance\napproval-request, approval-decision"]
KNOW["Knowledge\ndocument, runbook"]
LEARN["Learning\npattern, skill, feedback"]
end
subgraph governance["Layer 2 — Governance (who controls what)"]
direction LR
IDENT["Identity\nperson, agent, identity-provider"]
SEC["Secrets\nsecret, key, access-grant"]
POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
end
subgraph infra["Layer 1 — Infrastructure (the managed world)"]
direction LR
PHYS["Physical\nsite, machine, ups, sensor"]
COMP["Compute\nmachine, vm, container\n(lxc, docker)"]
NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"]
end
Block definition diagram (SysML BDD)
Uses Mermaid class diagram syntax following SysML BDD conventions:
«abstract»= abstract block (cannot be instantiated)<\|--= generalization (is-a)*--= composition (whole-part, lifecycle dependency)o--= aggregation (whole-part, independent lifecycle)--= association (typed link)- Multiplicity:
"1","0..1","1..*","0..*","*"
Infrastructure layer — compute, storage, network:
classDiagram
class ComputeEntity {
<<abstract>>
+state lifecycle
+attributes jsonb
}
class Machine {
+cpu_arch
+ram_gb
}
class VirtualMachine {
+vcpus
+memory_mb
+disk_gb
}
class Container {
<<abstract>>
+runtime
}
class LXC {
+pve_id
+rootfs
}
class DockerContainer {
+image
+compose_stack
}
class ProxmoxHost {
+pve_version
+cluster_member
}
class StandaloneServer {
+hypervisor
}
class Workstation {
+os
+user
}
class Appliance {
+vendor
+model
}
class Hypervisor {
+type
+version
}
ComputeEntity <|-- Machine
ComputeEntity <|-- VirtualMachine
ComputeEntity <|-- Container
Machine <|-- ProxmoxHost
Machine <|-- StandaloneServer
Machine <|-- Workstation
Machine <|-- Appliance
Container <|-- LXC
Container <|-- DockerContainer
Machine "1" *-- "0..1" Hypervisor : runs
Hypervisor "1" o-- "0..*" VirtualMachine : hosts
Hypervisor "1" o-- "0..*" Container : hosts
class StoragePool {
+type lvm, zfs, nfs
+capacity_gb
}
class Volume {
+name
+size_gb
}
class Mount {
+mount_point
+options
}
StoragePool "1" *-- "0..*" Volume : contains
ComputeEntity "1" o-- "0..*" Mount : has
Mount "0..*" --> "1" Volume : mounts
class NetworkInterface {
+mac
+ip
}
class Network {
<<abstract>>
}
class LAN { +subnet}
class Mesh { +provider}
class VLAN { +tag}
ComputeEntity "1" *-- "0..*" NetworkInterface : has
NetworkInterface "0..*" --> "1" Network : connects-to
Network <|-- LAN
Network <|-- Mesh
Network <|-- VLAN
Software + services layer:
classDiagram
class Service {
+port
+health_url
+risk_notes
}
class Application {
+version
+config
}
class ConfigRepo {
+url
+branch
}
class DeployPipeline {
+trigger
+target_path
}
class IngressRoute {
+pattern
+upstream
}
class Certificate {
+issuer
+expires
}
class DNSZone {
+zone
}
class DNSRecord {
+name
+record_type
+value
}
ComputeEntity "1" o-- "0..*" Service : provides
Service "1" *-- "0..*" Application : runs
Service "0..1" --> "0..1" ConfigRepo : configured-by
DeployPipeline "0..*" --> "1" Service : deploys-to
IngressRoute "0..*" --> "1" Service : routes-to
IngressRoute "0..*" --> "0..1" Certificate : secured-by
IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by
Service "0..*" --> "0..*" Service : depends-on
DNSZone "1" *-- "0..*" DNSRecord : contains
DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to
Cognition layer — operations + learning:
classDiagram
class Signal {
+kind
+severity
+state lifecycle
+evidence
+recommended_action
}
class Execution {
+status
+result
+duration_ms
+verified
}
class Feedback {
+outcome
+observation
+lesson
}
class Pattern {
+confidence
+evidence_count
+status
}
class Skill {
+procedure
+version
+success_rate
}
class Classification {
+risk
+route
+reasoning
}
class Approval {
+status
+ttl
}
class Document {
+title
+content
+source_path
}
class Runbook {
+steps
+risk_class
+verification
}
ComputeEntity "1" o-- "0..*" Signal : monitored-by
Signal "0..1" --> "0..*" Execution : triggers
Classification "1" --> "0..1" Execution : precedes
Execution "1" *-- "0..1" Feedback : produces
Feedback "0..*" --> "0..*" Pattern : contributes-to
Pattern "0..*" --> "0..1" Skill : informs
Skill "0..1" --> "0..*" Classification : guides
Execution "0..1" --> "0..1" Approval : requires
Agent "0..*" --> "0..*" Execution : performs
Person "0..1" --> "0..*" Approval : decides
Person "0..1" --> "0..*" Agent : owns
Entity "1" o-- "0..*" Document : documented-by
Entity "1" o-- "0..*" Runbook : procedure-for
Design notes — validated assumptions
Can VMs mount storage pools? Yes. In Proxmox, VMs have virtual disks on storage
pools (LVM, ZFS, NFS). LXCs have bind mounts and mount points. Both use the same
storage pools. The model reflects this: ComputeEntity (abstract) has Mount edges
to Volume, regardless of whether the compute entity is a VM, LXC, or machine.
Not every machine is a Proxmox host. The current homelab has: 2 Proxmox hosts, 2
workstations, 1 external VPS, 2 VMs, 19 LXCs. The model uses Machine as the base
type with specializations (ProxmoxHost, StandaloneServer, Workstation,
Appliance). Only ProxmoxHost runs PVE; StandaloneServer could run KVM/libvirt;
Workstation runs desktop OS + optionally Docker. A Hypervisor is software that
runs on a Machine and hosts VMs/containers — it's not always present (workstations
and appliances may not have one).
Docker containers are first-class compute entities. The OS itself runs in Docker
containers, and services like Jellyfin's MariaDB sidecar run in Docker within LXCs.
DockerContainer is a specialization of Container with image, compose_stack
attributes. This lets the OS model its own infrastructure.
Services run on any compute entity. The homelab has services on LXCs (caddy,
gitea), VMs (zimaos, haos), workstations (mac-mini will run the OS), and external
hosts (authentik on netbird-vps). The provides relationship is from
ComputeEntity (abstract), not from a specific compute type.
Documents and runbooks describe any entity. The old ER diagram had
Document → Service only. In practice, docs describe hosts, containers, network
infrastructure, storage, and investigations. The model uses Entity (the root
abstract type) for documented-by and procedure-for, so any entity can have docs
and runbooks.
Infrastructure lifecycle
stateDiagram-v2
[*] --> planned : operator creates entity
planned --> provisioning : IP reserved, storage chosen, doc stub
provisioning --> active : mesh joined, health-check answering, doc complete
active --> migrating : preflight + backup verified
migrating --> active : post-verify, caddy checked, mounts checked
active --> deprecated : replacement live or role retired
deprecated --> destroyed : backups verified, secrets revoked, ingress removed
destroyed --> [*] : archaeology entry recorded
note right of deprecated
Complete only when
zero inbound edges remain
(no depends-on, no routes-to)
end note
Signal lifecycle
stateDiagram-v2
[*] --> raised : scheduler probe or agent finding
raised --> acknowledged : agent or operator sees it
raised --> muted : operator suppresses (TTL)
acknowledged --> acting : actuator starts execution
acting --> resolved : action succeeded, verification passed
acting --> raised : action failed, re-escalated
raised --> resolved : condition cleared (auto-resolve)
muted --> raised : TTL expired
resolved --> [*]
Execution + learning lifecycle
stateDiagram-v2
[*] --> proposed : signal + recommended_action
proposed --> approved : operator approves (if gated)
proposed --> auto_approved : risk class allows auto-act
approved --> executing : actuator runs
auto_approved --> executing : actuator runs
executing --> verified : verification command passed
executing --> failed : execution or verification failed
executing --> timed_out : exceeded duration limit
failed --> rolled_back : rollback procedure executed
verified --> [*] : feedback recorded, pattern updated
failed --> [*] : feedback recorded, pattern updated
rolled_back --> [*] : feedback recorded, pattern updated
timed_out --> [*] : feedback recorded, pattern updated
Feedback / learning model — the cognition loop
flowchart TB
subgraph observe["Observe"]
SIG["Signal raised\n(service down, disk full, drift)"]
end
subgraph decide["Decide"]
CLASS["Classification\nrisk × blast radius × confidence"]
SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"]
CLASS --> SKILL_LOOKUP
end
subgraph act["Act"]
EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"]
VERIFY["Verification\ncheck if action succeeded"]
EXEC --> VERIFY
end
subgraph learn["Learn (feedback loop)"]
OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"]
FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"]
PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"]
SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"]
OUTCOME --> FEEDBACK
FEEDBACK --> PATTERN
PATTERN --> SKILL_REFINE
end
SIG --> CLASS
SKILL_LOOKUP --> EXEC
VERIFY --> OUTCOME
SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP
subgraph escalation["Escalation path"]
APPROVAL["Approval request\n→ Matrix ✅/❌"]
end
CLASS -- "needs approval" --> APPROVAL
APPROVAL -- "approved" --> EXEC
APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"]
Policy model — how risk attaches to entities
flowchart LR
subgraph ontology["Ontology (DB metadata)"]
ET["entity_types\nhost, service, lxc, vm, ..."]
RT["relationship_types\nhosts, provides, depends-on, ..."]
LD["lifecycle_defs\nstates + transitions"]
end
subgraph policy["Policy (DB records)"]
RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"]
RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"]
AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"]
end
subgraph instances["Instances (DB data)"]
ENT["entities\nhost:hubris, service:caddy, ..."]
REL["relationships\nhubris hosts lxc:apps"]
end
ET --> RULES
RC --> RULES
RULES --> ENT
AUTO --> ENT
ET --> ENT
RT --> REL
LD --> ENT
Target architecture
Container stack on mac-mini
graph TB
subgraph mac-mini["mac-mini — Docker host, always-on"]
subgraph services["Docker Compose — Go binaries"]
PG["PostgreSQL 16 + TimescaleDB\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata\n• metric_samples (hypertable)\n• audit_log (hypertable)\n• events (hypertable)\n• agent_activity (hypertable)"]
INF["Infisical\n(secrets manager)"]
HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"]
API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations,\nquery_metrics, get_trend,\nget_audit_trail, get_health_summary\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events,\n/metrics, /audit, /health\n\nPolicy enforcement + risk\nclassification + ledger\nAudit middleware + event emitter"]
SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB + metrics\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• Correlation ID propagation\n• SSH to hubris/strong"]
NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"]
end
DEPLOY["Gitea webhook →\ndocker compose build + up -d"]
end
HERMES -- MCP --> API
API --> PG
SCHED --> PG
SCHED -- "metrics + events + audit" --> PG
API -- "audit + events" --> PG
HERMES -- "agent_activity" --> PG
SCHED -- SSH --> HUBRIS
SCHED -- SSH --> STRONG
API -- "escalate" --> NOTIFIER
NOTIFIER -- "alerts + approvals" --> MATRIX
DEPLOY -- "rebuild" --> services
subgraph external["External"]
CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"]
APPS["apps/105 (fallback)"]
HUBRIS["hubris (PVE)"]
STRONG["strong (PVE)"]
GITEA["Gitea (LXC 104)"]
MATRIX["Matrix (LXC 118)"]
WS["Any workstation\nHermes remote → gateway"]
end
CADDY -- reverse_proxy --> API
GITEA -- webhook --> DEPLOY
WS -- "Hermes gateway" --> HERMES
Observability — data capture and query flow
flowchart TB
subgraph sources["Data sources"]
SCHED_PROBE["Scheduler probes\nhealth, disk, latency, drift"]
API_CALLS["API calls\nrequest/response metrics"]
AGENT_CALLS["Agent activity\nMCP tools, SSH, reasoning"]
EXEC["Executions\nSSH commands, verification"]
LEARN["Learning engine\npatterns, skills, confidence"]
DEPLOY_EVT["Deploy events\nwebhook, build, restart"]
end
subgraph capture["Capture layer (Go)"]
METRICS_W["metrics.go\n→ metric_samples"]
AUDIT_W["audit.go\n→ audit_log"]
EVENTS_W["events.go\n→ events table"]
AGENT_W["agent.go\n→ agent_activity"]
end
subgraph store["PostgreSQL + TimescaleDB"]
MS["metric_samples\n(hypertable, 90d raw,\n1y rollups via CAGGs)"]
AL["audit_log\n(hypertable, 1y retention)"]
EV["events\n(hypertable, 90d retention)"]
AA["agent_activity\n(hypertable, 90d retention)"]
end
subgraph query["Query layer (API)"]
REST["REST endpoints\n/metrics, /audit,\n/events, /health,\n/trends, /agent-activity"]
MCP["MCP tools\nquery_metrics, get_trend,\nget_audit_trail, get_event_timeline,\nget_agent_activity, get_health_summary"]
WS["WebSocket\n/api/v1/events\n(live stream)"]
end
subgraph consumers["Consumers"]
AGENT["Agent\nqueries trends, audits\nits own history"]
FUTURE["Future: Grafana,\ndashboards, notebooks,\nSIEM, compliance tools"]
end
SCHED_PROBE --> METRICS_W
API_CALLS --> METRICS_W
API_CALLS --> AUDIT_W
AGENT_CALLS --> AGENT_W
EXEC --> EVENTS_W
EXEC --> AUDIT_W
LEARN --> METRICS_W
LEARN --> EVENTS_W
DEPLOY_EVT --> EVENTS_W
METRICS_W --> MS
AUDIT_W --> AL
EVENTS_W --> EV
AGENT_W --> AA
MS --> REST
AL --> REST
EV --> REST
EV --> WS
AA --> REST
MS --> MCP
AL --> MCP
EV --> MCP
AA --> MCP
REST --> AGENT
MCP --> AGENT
REST --> FUTURE
WS --> FUTURE
OODA loop — with learning feedback
flowchart LR
OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"]
ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"]
DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"]
DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"]
APPROVE -- "approved" --> ACT
ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"]
LEARN -. "improves confidence" .-> DECIDE
LEARN --> OBSERVE
Deploy flow
flowchart LR
DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"]
GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"]
MACMINI -- "git pull" --> REPO["repo clone"]
MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"]
STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"]
DB-native configuration
The three YAML files — inventory.yaml, ontology.yaml, policy.yaml — become
seed manifests. They bootstrap the DB on first deploy. After that, the DB is the
runtime source of truth, editable via the API. A future frontend can edit all three
directly.
How it works
flowchart TB
subgraph seeds["Seed manifests (git-tracked, YAML)"]
ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"]
INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"]
POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"]
end
subgraph db["PostgreSQL (runtime source of truth)"]
META["entity_types table\nrelationship_types table\nlifecycle_defs table"]
INST["entities table\nrelationships table"]
POLDB["policies table\nrisk_classes table\nautonomy_settings table"]
end
INGEST["Seed ingest (on deploy)\nidempotent upsert"]
ONTO_YAML --> INGEST --> META
INV_YAML --> INGEST --> INST
POL_YAML --> INGEST --> POLDB
API_EDIT["API edits\n(POST/PUT/PATCH)"]
API_EDIT --> META
API_EDIT --> INST
API_EDIT --> POLDB
EXPORT["Export to YAML\n(for DR / version control)"]
META --> EXPORT
INST --> EXPORT
POLDB --> EXPORT
Why DB-native
- Querying — the agent can ask "what services depend on authentik?" as a graph query, not a YAML parse. Blast-radius walks are SQL, not file reads.
- Mutation — adding a service, updating a lifecycle state, changing a policy rule are DB transactions with audit trail, not file edits + git commits.
- Consistency — the ontology, inventory, and policy are always in sync (same DB, same transaction). No drift between what the YAML says and what the runtime sees.
- Future frontend — a UI can edit entities, relationships, and policies directly via the API. No need to generate/edit YAML files.
- Version control — seed YAML files are still git-tracked for bootstrap and DR. The API can export the current DB state back to YAML for commit.
Repo layout (Go project)
/ # repo root
├── docker-compose.yml # the OS stack definition
├── Makefile # build, test, deploy targets
├── go.mod # Go module definition
├── go.sum
├── cmd/ # binary entrypoints (one per service)
│ ├── api/ # Oikos API server
│ │ └── main.go
│ ├── scheduler/ # Observe + Act loop
│ │ └── main.go
│ └── notifier/ # Notification service
│ └── main.go
├── internal/ # private packages (not importable)
│ ├── db/ # database layer
│ │ ├── queries/ # sqlc SQL queries
│ │ ├── models.go # generated Go types
│ │ └── db.go # connection pool, migrations
│ ├── ontology/ # ontology types + meta-schema
│ │ ├── types.go # EntityType, RelationshipType, LifecycleDef
│ │ ├── graph.go # graph traversal (blast radius, dependencies)
│ │ └── ingest.go # YAML seed → DB ingest
│ ├── api/ # HTTP + MCP server
│ │ ├── server.go # Gin app setup
│ │ ├── routes/ # REST handlers
│ │ │ ├── hosts.go
│ │ │ ├── services.go
│ │ │ ├── signals.go
│ │ │ ├── approvals.go
│ │ │ ├── exec.go
│ │ │ └── knowledge.go
│ │ └── mcp.go # MCP protocol adapter (JSON-RPC over SSE)
│ ├── policy/ # risk classification + approval
│ │ ├── classify.go # risk × blast × confidence
│ │ ├── approve.go # approval request + grant lifecycle
│ │ └── autonomy.go # kill-switch, never-auto-act list
│ ├── scheduler/ # Observe stage
│ │ ├── probe.go # HTTP health, disk, drift
│ │ └── signal.go # raise/resolve signals in DB
│ ├── actuator/ # Act stage
│ │ ├── act.go # read signals, classify, execute or escalate
│ │ ├── execute.go # SSH execution + verification
│ │ └── guard.go # loop-guard, retry caps
│ ├── learning/ # feedback loop (the learning model)
│ │ ├── feedback.go # record outcome + lesson from execution
│ │ ├── pattern.go # extract/validate patterns from feedback
│ │ └── skill.go # create/refine skills from patterns
│ ├── notifier/ # notification abstraction
│ │ ├── notifier.go # interface
│ │ └── matrix.go # Matrix implementation
│ ├── observability/ # logging, metrics, audit, events (NEW)
│ │ ├── logging.go # slog structured logging setup
│ │ ├── metrics.go # metric recording (writes to metric_samples)
│ │ ├── audit.go # audit middleware (writes to audit_log)
│ │ ├── events.go # event emitter (writes to events table)
│ │ ├── agent.go # agent activity recording
│ │ └── correlation.go # correlation ID propagation (context-based)
│ └── config/ # config loading (env, files)
│ └── config.go
├── migrations/ # SQL migrations (golang-migrate format)
│ ├── 001_ontology.up.sql # meta-schema (entity_types, relationship_types, ...)
│ ├── 001_ontology.down.sql
│ ├── 002_instances.up.sql # entities, relationships
│ ├── 003_operations.up.sql # signals, approvals, executions, feedback
│ ├── 004_learning.up.sql # patterns, skills
│ ├── 005_policy.up.sql # policies, risk_classes, autonomy
│ └── 006_observability.up.sql # metrics, audit_log, events, agent_activity
├── seeds/ # YAML seed manifests (bootstrap + DR)
│ ├── ontology.yaml # entity types, relationship types, lifecycles
│ ├── inventory.yaml # entity instances (hosts, services, etc.)
│ └── policy.yaml # risk classes, approval rules, autonomy
├── docs/ # narrative docs (ingested into knowledge graph)
│ ├── containers/
│ ├── hosts/
│ ├── infrastructure/
│ └── investigations/
├── hermes/ # Hermes agent config + skills
│ ├── config.yaml
│ ├── SOUL.md
│ └── skills/
│ └── homelab-ops/
│ └── SKILL.md
├── compose/ # Docker build contexts
│ ├── api/Dockerfile
│ ├── scheduler/Dockerfile
│ ├── hermes/Dockerfile
│ └── postgres/init.sql
└── scripts/ # utility scripts
├── migrate-sops.sh # one-time SOPS → Infisical migration
└── import-legacy.sh # import existing signals/ledger JSONL
Database schema
The schema is the ontology made concrete. Five migration groups, each adding a layer.
Migration 1: Ontology meta-schema
-- The meta-graph: defines what entity types and relationship types can exist.
-- This IS the ontology, stored in the DB, editable via API.
CREATE TABLE entity_types (
name TEXT PRIMARY KEY, -- 'host', 'service', 'signal', 'pattern'
domain TEXT NOT NULL, -- 'physical', 'compute', 'network', ...
layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition'
description TEXT,
lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle)
attribute_schema JSONB, -- JSON Schema for validating entity.attributes
status TEXT NOT NULL DEFAULT 'active', -- 'active', 'deprecated' (no hard delete while instances exist)
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE TABLE relationship_types (
name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on'
inverse TEXT, -- 'runs-on', 'provided-by'
source_type TEXT REFERENCES entity_types(name),
target_type TEXT REFERENCES entity_types(name),
cardinality TEXT NOT NULL, -- 'one-to-one', 'one-to-many', 'many-to-many'
description TEXT,
created_at TIMESTAMPTZ DEFAULT now()
);
CREATE TABLE lifecycle_defs (
id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', 'execution', ...
states TEXT[] NOT NULL, -- ordered states
default_state TEXT NOT NULL,
transitions JSONB NOT NULL, -- {"from": {"to": {"requires": [...]}}}
created_at TIMESTAMPTZ DEFAULT now()
);
Migration 2: Entity instances (the inventory graph)
-- Entity instances — the actual hosts, services, signals, patterns, etc.
-- This replaces inventory.yaml as the runtime source of truth.
CREATE TABLE entities (
id TEXT PRIMARY KEY, -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001'
type TEXT NOT NULL REFERENCES entity_types(name),
name TEXT NOT NULL, -- 'hubris', 'caddy', 'disk-threshold'
state TEXT, -- lifecycle state (e.g. 'active', 'raised')
attributes JSONB NOT NULL DEFAULT '{}', -- type-specific data (IP, mesh addr, port, ...)
parent_id TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host)
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_entities_type ON entities(type);
CREATE INDEX idx_entities_state ON entities(state);
CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes);
-- Relationship instances — the typed edges of the graph
CREATE TABLE relationships (
source_id TEXT NOT NULL REFERENCES entities(id),
target_id TEXT NOT NULL REFERENCES entities(id),
type TEXT NOT NULL REFERENCES relationship_types(name),
attributes JSONB,
created_at TIMESTAMPTZ DEFAULT now(),
PRIMARY KEY (source_id, target_id, type)
);
CREATE INDEX idx_rel_source ON relationships(source_id);
CREATE INDEX idx_rel_target ON relationships(target_id);
CREATE INDEX idx_rel_type ON relationships(type);
-- Recursive graph traversal function (blast radius, dependency chains)
CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3)
RETURNS TABLE(entity_id TEXT, depth INT) AS $$
WITH RECURSIVE walk AS (
SELECT start_id::TEXT AS entity_id, 0::INT AS depth
UNION
SELECT r.target_id::TEXT, w.depth + 1
FROM relationships r
JOIN walk w ON r.source_id = w.entity_id
WHERE w.depth < max_depth
)
SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
$$ LANGUAGE sql STABLE;
Migration 3: Operations (signals, approvals, ledger, state)
-- Signals — now entities in the graph, with a dedicated table for indexed querying
-- (the entity row is the canonical record; this table is a fast lookup)
CREATE TABLE signals (
entity_id TEXT PRIMARY KEY REFERENCES entities(id),
kind TEXT NOT NULL,
severity TEXT NOT NULL, -- info, warning, critical
target_entity_id TEXT REFERENCES entities(id), -- the infrastructure entity this is about
evidence TEXT,
likely_cause TEXT,
recommended_action JSONB,
verification TEXT,
state TEXT NOT NULL DEFAULT 'raised',
mute_until TIMESTAMPTZ,
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_signals_state ON signals(state);
CREATE INDEX idx_signals_target ON signals(target_entity_id);
CREATE INDEX idx_signals_severity ON signals(severity);
-- Approvals
CREATE TABLE approvals (
id TEXT PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id), -- entity to act on
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'pending', -- pending, approved, denied, expired
ttl INTERVAL NOT NULL DEFAULT '1 hour',
decided_at TIMESTAMPTZ,
decided_by TEXT REFERENCES entities(id), -- person entity
confirmation_phrase TEXT
);
-- Change ledger (high-level record, links to execution for detail)
CREATE TABLE ledger_entries (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id),
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
result TEXT, -- ok, failed, escalated
approval_id TEXT REFERENCES approvals(id),
execution_id INTEGER, -- FK to executions (migration 4)
agent_id TEXT REFERENCES entities(id),
notes TEXT
);
-- State snapshots (replaces oikos/state.json)
CREATE TABLE state_snapshots (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
entity_id TEXT REFERENCES entities(id),
health TEXT, -- healthy, degraded, down, unknown
data JSONB
);
Migration 4: Learning model (executions, feedback, patterns, skills)
-- Executions — detailed record of each action the OS performs
CREATE TABLE executions (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
signal_entity_id TEXT REFERENCES entities(id), -- signal that triggered this
target_entity_id TEXT REFERENCES entities(id), -- entity acted upon
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
approval_id TEXT REFERENCES approvals(id),
agent_id TEXT REFERENCES entities(id), -- who/what executed
skill_id TEXT REFERENCES entities(id), -- skill used (if any)
status TEXT NOT NULL DEFAULT 'queued', -- queued, running, completed, failed, timed-out
result JSONB, -- detailed result data
duration_ms INTEGER,
verified BOOLEAN DEFAULT false,
started_at TIMESTAMPTZ,
completed_at TIMESTAMPTZ
);
CREATE INDEX idx_exec_target ON executions(target_entity_id);
CREATE INDEX idx_exec_status ON executions(status);
CREATE INDEX idx_exec_action ON executions(action);
-- Feedback — what was learned from an execution
CREATE TABLE feedback (
id SERIAL PRIMARY KEY,
execution_id INTEGER REFERENCES executions(id),
ts TIMESTAMPTZ DEFAULT now(),
outcome TEXT NOT NULL, -- success, failure, partial, unexpected
observation TEXT, -- what happened vs what was expected
lesson TEXT, -- extractable lesson
unexpected_side_effects TEXT[],
tags TEXT[]
);
CREATE INDEX idx_feedback_execution ON feedback(execution_id);
CREATE INDEX idx_feedback_outcome ON feedback(outcome);
-- Patterns — generalized rules extracted from accumulated feedback
CREATE TABLE patterns (
id TEXT PRIMARY KEY, -- 'pat-2026-07-06-001'
ts TIMESTAMPTZ DEFAULT now(),
entity_type TEXT REFERENCES entity_types(name), -- applies to this type
action TEXT NOT NULL, -- 'restart', 'deploy', etc.
pattern TEXT NOT NULL, -- 'service X recovers within 30s after restart'
confidence REAL DEFAULT 0.5, -- 0.0 to 1.0
evidence_count INTEGER DEFAULT 1, -- how many executions support this
success_count INTEGER DEFAULT 0,
failure_count INTEGER DEFAULT 0,
status TEXT DEFAULT 'hypothesized', -- hypothesized, validated, active, deprecated
last_validated_at TIMESTAMPTZ
);
CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action);
CREATE INDEX idx_patterns_status ON patterns(status);
-- Skills — codified procedures refined through feedback
CREATE TABLE skills (
id TEXT PRIMARY KEY, -- 'skill-restart-service', 'skill-deploy-lxc'
name TEXT NOT NULL,
ts TIMESTAMPTZ DEFAULT now(),
procedure TEXT NOT NULL, -- the codified steps (markdown or structured)
applies_to TEXT REFERENCES entity_types(name),
pattern_ids TEXT[], -- patterns that inform this skill
status TEXT DEFAULT 'drafted', -- drafted, tested, active, refined, deprecated
version INTEGER DEFAULT 1,
success_rate REAL, -- rolling success rate
last_used_at TIMESTAMPTZ
);
CREATE INDEX idx_skills_type ON skills(applies_to);
CREATE INDEX idx_skills_status ON skills(status);
Migration 5: Policy (DB-native risk + approval rules)
-- Risk classes — the four-level safety model
CREATE TABLE risk_classes (
name TEXT PRIMARY KEY, -- 'read_only', 'reversible_low', etc.
description TEXT,
approval_required TEXT NOT NULL DEFAULT 'none', -- none, operator, operator_confirmed
ledger BOOLEAN DEFAULT false,
autonomy_allowed BOOLEAN DEFAULT false -- can agent auto-act at this risk level?
);
-- Approval rules — entity_type + action → risk_class + requirements
CREATE TABLE approval_rules (
id SERIAL PRIMARY KEY,
entity_type TEXT REFERENCES entity_types(name), -- applies to this entity type
action TEXT NOT NULL, -- 'restart', 'deploy', 'destroy'
risk_class TEXT NOT NULL REFERENCES risk_classes(name),
autonomy_level TEXT NOT NULL DEFAULT 'auto', -- 'auto', 'escalate', 'never'
scope_entity TEXT REFERENCES entities(id), -- optional: specific entity only
created_at TIMESTAMPTZ DEFAULT now(),
updated_at TIMESTAMPTZ DEFAULT now(),
UNIQUE(entity_type, action)
);
-- Autonomy settings — global kill-switch + per-entity overrides
CREATE TABLE autonomy_settings (
key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.caddy'
value TEXT NOT NULL, -- 'off', 'reversible_low', 'true', 'false'
updated_at TIMESTAMPTZ DEFAULT now()
);
Migration 6: Observability (metrics, audit log, event log)
Uses TimescaleDB (PostgreSQL extension) for time-series data. Hypertables auto-partition by time, continuous aggregates provide rollups, and retention policies auto-drop old data. No separate database needed — everything stays in Postgres.
-- Enable TimescaleDB
CREATE EXTENSION IF NOT EXISTS timescaledb;
-- ─── Time-series metrics ──────────────────────────────────────────────
-- Generic metric store. Every probe, health check, and system measurement
-- writes here. Designed for high insert volume, time-range queries, and
-- continuous-aggregate rollups.
CREATE TABLE metric_samples (
ts TIMESTAMPTZ NOT NULL,
entity_id TEXT NOT NULL, -- which entity this metric is about
metric TEXT NOT NULL, -- 'health', 'disk_usage_pct', 'probe_latency_ms',
-- 'api_p99_ms', 'goroutines', 'db_connections',
-- 'pattern_confidence', 'skill_success_rate', ...
value DOUBLE PRECISION NOT NULL,
tags JSONB DEFAULT '{}'::JSONB -- arbitrary key-value labels:
-- {probe: "http", target: "192.168.8.121"},
-- {host: "hubris", mount: "/mnt/library"}, ...
);
-- Hypertable: partition by time, 1-week chunks
SELECT create_hypertable('metric_samples', 'ts', chunk_time_interval => INTERVAL '7 days');
-- Indexes for common query patterns
CREATE INDEX idx_metrics_entity_ts ON metric_samples(entity_id, ts DESC);
CREATE INDEX idx_metrics_metric_ts ON metric_samples(metric, ts DESC);
CREATE INDEX idx_metrics_tags ON metric_samples USING GIN(tags);
-- Retention: drop raw metrics older than 90 days (continuous aggregates keep rollups)
SELECT add_retention_policy('metric_samples', INTERVAL '90 days');
-- Continuous aggregate: 1-hour rollups (mean, min, max, count)
CREATE MATERIALIZED VIEW metric_rollups_1h
WITH (timescaledb.continuous) AS
SELECT
time_bucket('1 hour', ts) AS bucket,
entity_id,
metric,
avg(value) AS avg_value,
min(value) AS min_value,
max(value) AS max_value,
count(*) AS sample_count,
(array_agg(tags))[1] AS representative_tags
FROM metric_samples
GROUP BY bucket, entity_id, metric;
-- Refresh policy: refresh every 1 hour, keep 1 year of rollups
SELECT add_continuous_aggregate_policy('metric_rollups_1h',
start_offset => INTERVAL '2 hours',
end_offset => INTERVAL '5 minutes',
schedule_interval => INTERVAL '1 hour');
CREATE MATERIALIZED VIEW metric_rollups_1d
WITH (timescaledb.continuous) AS
SELECT
time_bucket('1 day', ts) AS bucket,
entity_id,
metric,
avg(value) AS avg_value,
min(value) AS min_value,
max(value) AS max_value,
count(*) AS sample_count
FROM metric_samples
GROUP BY bucket, entity_id, metric;
SELECT add_continuous_aggregate_policy('metric_rollups_1d',
start_offset => INTERVAL '2 days',
end_offset => INTERVAL '1 hour',
schedule_interval => INTERVAL '1 day');
-- ─── Audit log ────────────────────────────────────────────────────────
-- Every mutating action — by the OS, by an agent, or by an operator —
-- gets an immutable audit entry. This is the "who did what when" trail
-- that the ledger doesn't fully capture (ledger records OS decisions and
-- executions; audit captures ALL API calls including reads-that-matter
-- and operator interactions).
CREATE TABLE audit_log (
id BIGSERIAL PRIMARY KEY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
actor_type TEXT NOT NULL, -- 'agent', 'operator', 'system', 'scheduler'
actor_id TEXT, -- entity ID of the actor (agent entity, person entity)
action TEXT NOT NULL, -- 'api.call', 'entity.create', 'policy.update',
-- 'approval.decide', 'exec.request', 'deploy.trigger'
entity_id TEXT, -- entity affected (if any)
method TEXT, -- 'GET', 'POST', 'PATCH', 'DELETE', 'MCP', 'SSH'
path TEXT, -- API path or MCP tool name or SSH command
status_code INTEGER, -- HTTP status or 0 for non-HTTP
detail JSONB DEFAULT '{}'::JSONB, -- request body, response summary, extra context
source_ip TEXT, -- where the call came from
correlation_id TEXT -- links to execution_id / signal_id for tracing
);
SELECT create_hypertable('audit_log', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_audit_actor ON audit_log(actor_type, actor_id, ts DESC);
CREATE INDEX idx_audit_entity ON audit_log(entity_id, ts DESC);
CREATE INDEX idx_audit_action ON audit_log(action, ts DESC);
CREATE INDEX idx_audit_correlation ON audit_log(correlation_id);
-- Retention: keep 1 year of audit logs
SELECT add_retention_policy('audit_log', INTERVAL '365 days');
-- ─── Event log ────────────────────────────────────────────────────────
-- Structured event stream — the "news feed" of the OS. Every significant
-- state change is an event: signal raised/resolved, execution started/completed,
-- approval requested/granted, deploy triggered, pattern validated, skill refined,
-- config changed, entity lifecycle transition. The WebSocket /api/v1/events
-- streams from this table; agents can also query it historically.
CREATE TABLE events (
id BIGSERIAL PRIMARY KEY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
type TEXT NOT NULL, -- 'signal.raised', 'signal.resolved',
-- 'execution.started', 'execution.completed',
-- 'approval.requested', 'approval.decided',
-- 'deploy.triggered', 'deploy.completed',
-- 'pattern.validated', 'skill.refined',
-- 'entity.created', 'entity.state_changed',
-- 'policy.changed', 'config.changed'
entity_id TEXT, -- primary entity involved
severity TEXT DEFAULT 'info', -- info, warning, critical
source TEXT NOT NULL, -- 'scheduler', 'actuator', 'api', 'hermes',
-- 'deploy', 'notifier', 'learning'
data JSONB DEFAULT '{}'::JSONB, -- event-specific payload
correlation_id TEXT -- links to signal/exec/approval for tracing
);
SELECT create_hypertable('events', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_events_type_ts ON events(type, ts DESC);
CREATE INDEX idx_events_entity_ts ON events(entity_id, ts DESC);
CREATE INDEX idx_events_severity_ts ON events(severity, ts DESC);
CREATE INDEX idx_events_correlation ON events(correlation_id);
-- Retention: keep 90 days of events (signals/ledger have their own tables
-- for permanent records; events are the transient feed)
SELECT add_retention_policy('events', INTERVAL '90 days');
-- ─── Agent activity log ──────────────────────────────────────────────
-- Records what the agent (Hermes) does: tool calls, reasoning, decisions,
-- token usage, latency. This is for agent behavior auditing and trend
-- analysis ("is the agent getting more efficient?").
CREATE TABLE agent_activity (
id BIGSERIAL PRIMARY KEY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
agent_id TEXT NOT NULL, -- entity ID of the agent
session_id TEXT, -- Hermes session ID
activity_type TEXT NOT NULL, -- 'tool_call', 'reasoning', 'decision',
-- 'mcp_query', 'ssh_command', 'escalation'
tool_name TEXT, -- MCP tool or CLI command called
entity_id TEXT, -- entity acted upon (if any)
input_summary TEXT, -- truncated input (first 500 chars)
output_summary TEXT, -- truncated output (first 500 chars)
duration_ms INTEGER,
token_count INTEGER, -- LLM tokens consumed (if applicable)
success BOOLEAN,
correlation_id TEXT
);
SELECT create_hypertable('agent_activity', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_agent_activity_agent_ts ON agent_activity(agent_id, ts DESC);
CREATE INDEX idx_agent_activity_type_ts ON agent_activity(activity_type, ts DESC);
CREATE INDEX idx_agent_activity_entity ON agent_activity(entity_id, ts DESC);
CREATE INDEX idx_agent_activity_correlation ON agent_activity(correlation_id);
-- Retention: keep 90 days of agent activity
SELECT add_retention_policy('agent_activity', INTERVAL '90 days');
What gets captured where:
| Data source | Table | Retention | Purpose |
|---|---|---|---|
| Scheduler probes (HTTP health, disk, drift) | metric_samples |
90 days raw, 1 year rollups | Trend analysis, anomaly detection |
| API response latencies | metric_samples |
same | Performance monitoring |
| Go runtime metrics (goroutines, mem, GC) | metric_samples |
same | OS self-monitoring |
| Learning model metrics (pattern confidence, skill success rate) | metric_samples |
same | Learning trend tracking |
| All mutating API calls | audit_log |
1 year | Auditing, compliance, forensics |
| Operator actions (approval decisions, entity edits, policy changes) | audit_log |
1 year | Operator accountability |
| Agent (Hermes) tool calls + reasoning | agent_activity |
90 days | Agent behavior auditing, efficiency tracking |
| Signal/exec/approval/deploy state changes | events |
90 days | Event stream (WebSocket), timeline reconstruction |
| Signals (permanent record) | signals |
no expiry | Signal lifecycle tracking |
| Ledger entries (permanent record) | ledger_entries |
no expiry | Change history |
| Executions (permanent record) | executions |
no expiry | Execution audit trail |
Workstreams
1. Ontology definition + seed manifests (seeds/, internal/ontology/)
Before any code, finalize the ontology. The current ontology.yaml has 8 domains
and 14 relationship types. The new ontology adds:
- Layer 3 entities (cognition): signal, execution, feedback, pattern, skill, approval, classification, document, runbook
- New relationships:
triggers,produces,contributes-to,informs,guides,precedes,performs,procedure-for,learned-from - Lifecycle definitions for operational entities (signals, executions, patterns, skills) — not just infrastructure
Deliverables:
seeds/ontology.yaml— entity types, relationship types, lifecycle definitions (seeded intoentity_types,relationship_types,lifecycle_defson deploy)seeds/inventory.yaml— adapted from currentinventory.yaml(seeded intoentities+relationships)seeds/policy.yaml— adapted from currentpolicy.yaml(seeded intorisk_classes,approval_rules,autonomy_settings)internal/ontology/ingest.go— idempotent seed → DB ingest
2. Database layer + migrations (migrations/, internal/db/)
- Write the 5 migrations above
- Set up sqlc for type-safe Go database access
- Connection pool, migration runner
- Graph traversal queries (blast radius, dependency chains, knowledge lookup)
- Import script for existing
signals/*.jsonl,ledger/*.jsonl→ DB
3. Unified API server — Go (cmd/api/, internal/api/)
One Go binary (Gin web framework) exposing REST + MCP from the same codebase.
MCP interface (internal/api/mcp.go):
- JSON-RPC over SSE, compatible with Hermes MCP client
- Tools:
get_host,list_services,search_knowledge,get_entity,get_relations,get_blast_radius,get_signal_history,get_ledger,get_patterns,get_skills,get_state_snapshot - Observability tools:
query_metrics(time-series for an entity/metric),get_trend(trend analysis with slope + anomaly flags),get_audit_trail(who-did-what for an entity or actor),get_event_timeline(structured event feed for an entity or time range),get_agent_activity(what the agent did),get_health_summary(current health across the fleet with trend indicators) - All read from PostgreSQL
REST interface (internal/api/routes/):
GET /api/v1/entities— list entities (filter by type, state, domain)GET /api/v1/entities/{id}— entity detail + relationshipsPOST /api/v1/entities— create entity (creates inventory entry)PATCH /api/v1/entities/{id}— update entity (state transition, attributes)GET /api/v1/signals— list signals (filter by state, severity, entity)POST /api/v1/signals/{id}/ack— acknowledgePOST /api/v1/signals/{id}/resolve— resolveGET /api/v1/approvals— pending approvalsPOST /api/v1/approvals/{id}/decide— approve/deny (Authentik-gated)POST /api/v1/exec— gated execution (classify → check approval → execute → feedback)GET /api/v1/patterns— list patterns (filter by entity_type, action, status)GET /api/v1/skills— list skills (filter by applies_to, status)GET /api/v1/knowledge/{entity_id}— knowledge graph queryGET /api/v1/knowledge/search?q=...— search knowledge graphGET /api/v1/ontology— list entity types, relationship types, lifecyclesPOST /api/v1/ontology/entity_types— create entity type (extend the schema)WS /api/v1/events— real-time stream (signals, approvals, executions, feedback)- Observability routes (
internal/api/routes/observability.go):GET /api/v1/metrics— query time-series:?entity_id=&metric=&from=&to=&interval=- Returns raw samples or rollups (auto-selects 1h/1d aggregates based on range)
?rollup=1h|1d|rawto force a specific resolution- Supports multiple metrics:
?metric=disk_usage_pct&metric=probe_latency_ms
GET /api/v1/metrics/{entity_id}/{metric}— single metric for one entity?from=2026-07-01T00:00:00Z&to=2026-07-06T00:00:00Z&rollup=1h- Returns:
{entity_id, metric, samples: [{ts, avg, min, max, count}], trend: {slope, direction, anomaly}}
GET /api/v1/trends/{entity_id}— trend analysis for all metrics on an entity- Returns slope (improving/degrading/stable), recent anomalies, forecast (simple linear)
GET /api/v1/audit— audit log:?actor_type=&actor_id=&entity_id=&action=&from=&to=- Paginated, ordered by ts DESC
?correlation_id=to trace a full execution chain
GET /api/v1/events— historical events:?type=&entity_id=&severity=&from=&to=- Same data as the WebSocket stream, but queryable historically
GET /api/v1/agent-activity— agent behavior log:?agent_id=&activity_type=&entity_id=&from=&to=- Includes token usage, latency, success/failure per tool call
GET /api/v1/health— fleet health summary with trend indicators- Returns:
{entities: [{id, type, health, trend, last_probe}], summary: {healthy, degraded, down, unknown}}
- Returns:
GET /api/v1/export— export current DB state as YAML (for DR / version control)
Policy enforcement (internal/policy/classify.go):
- Every mutating endpoint classifies the action via the policy DB
- Risk class → approval check → autonomy check
- All mutations write to the ledger automatically
Auth:
- MCP interface: no auth (internal, container-to-container)
- REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints
- Internal: shared secret (Docker network)
4. Scheduler + Actuator — Go (cmd/scheduler/, internal/scheduler/, internal/actuator/)
Scheduler (Observe) — Go service with goroutines for concurrent probes:
- HTTP health probes (concurrent, with timeouts)
- Disk usage probes (SSH to hubris/strong)
- Drift detection (inventory vs live state)
- Writes signals + state snapshots to DB
- Runs on a 10-min ticker
Actuator (Act) — Go service, the control loop:
- Reads open signals with
recommended_action - For each: classify via
internal/policy/classify.go- auto-act: look up skill for (entity_type, action) → follow procedure → execute via SSH → verify → record execution → generate feedback → update patterns
- escalate: create approval request → notify via Matrix → acknowledge signal
- Loop-guard: check execution history per (entity, action) to cap auto-retries
(
SELECT ... FOR UPDATE SKIP LOCKEDfor concurrency safety) - Autonomy kill-switch: check
autonomy_settingstable
5. Learning engine — Go (internal/learning/)
The feedback loop that makes the agent improve over time.
Feedback recording (internal/learning/feedback.go):
- After every execution, evaluate the outcome:
- Did the verification command pass? → success
- Did it fail? → failure
- Did it partially work? → partial
- Did something unexpected happen? → unexpected
- Record a feedback entry with: outcome, observation (what happened vs expected), lesson (extractable insight), unexpected_side_effects
Pattern extraction (internal/learning/pattern.go):
- Periodically scan accumulated feedback for (entity_type, action) pairs
- When N+ executions share a similar outcome, extract a pattern:
- "Restarting service:X typically takes 15s and succeeds"
- "Deploying to LXC:Y via webhook has 30% failure rate, retry helps"
- Patterns start as
hypothesized, move tovalidatedafter enough evidence, thenactive(used by the decision classifier) - Confidence score = success_count / evidence_count, adjusted by recency
Skill management (internal/learning/skill.go):
- When a pattern reaches
activestatus with confidence > 0.7, create or refine a skill for that (entity_type, action) pair - Skills codify the best-known procedure (what steps to take, what to verify, expected duration, known failure modes)
- Skills are versioned — each refinement increments the version
- The actuator looks up skills before executing: if a skill exists, follow it; if not, use the default procedure and generate feedback for future pattern extraction
How the classifier uses learning (internal/policy/classify.go):
- Confidence scoring now checks patterns + skills, not just raw ledger history:
- If a pattern exists for (entity_type, action) with high confidence → boost auto-act confidence
- If patterns show frequent failures → lower confidence, escalate
- If a skill exists → higher confidence (proven procedure available)
- This is the closed loop: execution → feedback → pattern → skill → classification → execution (better informed each time)
6. Knowledge graph ingestion (internal/ontology/ingest.go)
After the ontology is defined and the DB schema is in place:
- On deploy, walk
docs/directory - Parse each markdown file:
- Extract frontmatter for metadata (entity type, tags, relations)
- Infer entity relationships from path conventions:
docs/containers/105-apps.md→ relationship toentity:lxc:apps - Extract cross-references (markdown links) → relationships
- Create knowledge entities in the
entitiestable (type =document,runbook,investigation, etc.) withdocuments/procedure-foredges to infrastructure entities - Idempotent — safe to re-run on every deploy
Agent access:
- MCP tool
search_knowledge(query)— full-text search on knowledge entities - MCP tool
get_entity_knowledge(entity_id)— all docs related to an entity - MCP tool
get_relations(entity_id)— graph traversal (blast radius, dependencies)
7. Hermes agent container (compose/hermes/)
- Hermes Agent runtime in a Docker container, gateway mode
- Config:
hermes/config.yaml(providers, models, gateway port) - Persona:
hermes/SOUL.md(homelab-specific) - Skills:
hermes/skills/homelab-ops/SKILL.md— how to use the Oikos API, classify actions, request approvals, query the knowledge graph - Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid)
- Gateway port 8092 — workstations connect remotely
- Hermes data volume for persistent state
8. Infisical secrets migration
Same as rev 1:
- Stand up Infisical in the Docker stack
- Migrate SOPS secrets (one-time decrypt + import)
- Wire all Go services to Infisical via machine identity
- Retire SOPS + age keys
9. Notifier — Go (cmd/notifier/, internal/notifier/)
Interface (Go):
type Notifier interface {
SendAlert(ctx context.Context, signal Signal) error
SendApprovalRequest(ctx context.Context, approval Approval) error
ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error)
}
Matrix implementation:
- Sends alerts to
@dtoro:avisperovia Synapse (LXC 118) - Approval requests as messages with ✅/❌ reactions
- Listens for reactions to record decisions
- Pluggable — future implementations (webhook, email) register via config
10. Docker build + deploy pipeline
- Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch)
docker compose buildfrom repo root- Gitea webhook on push to
main→ deploy script on mac-mini - Deploy script:
git pull && go build ./... && docker compose build && docker compose up -d - Seed ingest runs as part of the API startup (idempotent)
- Health checks on each service
11. Ingress re-point (Caddy)
dtoro/caddy-conf: pointmcp.hubris.network+oikos.hubris.network→ mac-mini mesh IP :8090 (API)- Future:
hermes.hubris.network→ mac-mini:8092 - DNS and public URLs unchanged
12. mac-mini host setup
- New directory:
~/oikos-os/— repo clone +docker composeworking dir - Existing
/opt/homelab-context/stays untouched until cleanup phase - Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys
- Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone
13. Decommission apps/105 (deferred)
Keep apps/105 running as fallback. Cutover checklist when ready:
- Verify Docker OS serves all traffic
systemctl disable --nowOikos services on apps/105- Remove old checkouts
- Update Caddy backends exclusively to mac-mini
- Remove/retarget Gitea webhooks
14. Observability + data capture — Go (internal/observability/)
The OS captures three classes of data, all in PostgreSQL (TimescaleDB for time-series, regular tables for audit/events):
A. Metrics (time-series) — internal/observability/metrics.go:
- Infrastructure probes: every scheduler probe writes metrics:
health(0=down, 1=degraded, 2=healthy) per servicedisk_usage_pctper mount point per hostprobe_latency_msper probe targetdrift_countper drift check
- OS self-metrics: the API server exposes a
/metricsendpoint and also writes to the DB:api_request_count(per route, per status code)api_latency_ms(p50, p99)db_connections_activedb_query_duration_msgoroutines(Go runtime)gc_pause_msmemory_alloc_mb
- Learning metrics: the learning engine writes:
pattern_confidenceper pattern IDskill_success_rateper skill IDauto_act_countvsescalation_count(autonomy ratio over time)execution_duration_msper (entity_type, action)
- Agent metrics: Hermes activity is logged:
agent_token_countper sessionagent_tool_call_countper sessionagent_decision_latency_ms
- All metrics are written to
metric_samplesas hypertable with 1h and 1d continuous aggregates. Raw data retained 90 days, rollups 1 year.
B. Audit log (immutable) — internal/observability/audit.go:
- Gin middleware: every mutating API call (POST/PATCH/DELETE) writes to
audit_logwith actor, action, entity, method, path, status, detail, source_ip - MCP tool calls also audited (actor_type='agent')
- SSH commands by the actuator audited (actor_type='system', method='SSH')
- Policy mutations get special audit entries with a hash of the before/after state
correlation_idpropagated through the call chain (API → actuator → SSH → result) so a full execution chain can be reconstructed: "signal → classification → execution → SSH command → verification → feedback" all linked by correlation_id
C. Event stream — internal/observability/events.go:
- Every significant state change emits an event to the
eventstable:- Signal lifecycle:
signal.raised,signal.acknowledged,signal.acting,signal.resolved,signal.muted - Execution:
execution.started,execution.completed,execution.failed - Approval:
approval.requested,approval.decided - Deploy:
deploy.triggered,deploy.completed - Learning:
pattern.validated,skill.refined - Entity:
entity.created,entity.state_changed - Policy:
policy.changed
- Signal lifecycle:
- The WebSocket
/api/v1/eventsstreams from this table (new events pushed to subscribers, historical events queryable via REST) - Events carry
correlation_idfor end-to-end tracing
D. Agent activity — internal/observability/agent.go:
- Hermes MCP tool calls logged to
agent_activitywith tool name, entity acted upon, duration, token count, success/failure - SSH commands by the agent logged separately
- Reasoning/decision audit: when the agent makes a classification decision, the reasoning is recorded (input, classification result, route, why)
- This is the "what is the agent doing and is it getting better?" dataset
E. Structured logging — internal/observability/logging.go:
- All services use Go's
slog(structured JSON logging) - Every log line has:
ts,level,service,msg,correlation_id(when applicable),entity_id(when applicable) - Logs go to stdout (Docker captures them,
docker compose logsfor access) - Debug mode:
debug=trueenv var enables verbose probe payloads, SQL queries, classification reasoning in logs
F. Data availability for agents:
- The agent can query its own history: "what actions have I taken on service:caddy in the last 30 days, and what were the outcomes?"
- The agent can see trends: "is disk usage on hubris trending upward?"
- The agent can audit: "who changed the policy for service:caddy and when?"
- The agent can self-assess: "am I getting more efficient? (token usage trend)"
- All through MCP tools:
query_metrics,get_trend,get_audit_trail,get_event_timeline,get_agent_activity,get_health_summary
G. Future visualization plug-in points:
- The
/api/v1/metricsREST endpoint returns JSON time-series — any tool (Grafana, custom dashboard, notebook) can consume it - The data model is compatible with Grafana's PostgreSQL data source (time column + metric name + value + tags as labels)
- The event stream (
/api/v1/events+ WebSocket) can feed a live dashboard - The audit log can feed a SIEM or compliance tool
- No UI built now — the APIs are the contract; visualization plugs in later
Phasing
Phase 0 — Ontology design (no code):
- Finalize entity types, relationship types, lifecycles
- Write seed manifests (
seeds/ontology.yaml,seeds/inventory.yaml,seeds/policy.yaml) - Review diagrams with operator
Phase 1 — Foundation (Go + DB):
- Go module setup, project structure
- PostgreSQL + TimescaleDB setup
- Migrations 1-6 (ontology, instances, operations, learning, policy, observability)
- Seed ingest pipeline (YAML → DB)
- sqlc queries for core operations
- Structured logging setup (slog)
- Import existing signals/ledger data
Phase 2 — API (Go):
- Gin server with REST routes
- MCP protocol adapter (including observability tools)
- Policy enforcement middleware
- Audit middleware (every mutating call → audit_log)
- Event emitter (every state change → events table)
- Knowledge graph ingestion from docs/
- Observability routes (metrics, trends, audit, events, health)
Phase 3 — Control loop (Go):
- Scheduler (Observe) — probes, signals, state snapshots, metric recording
- Actuator (Act) — classify, execute, verify, correlation ID propagation
- Learning engine — feedback, patterns, skills, learning metrics
- Circuit breaker per target host
Phase 4 — Agent (Hermes container):
- Hermes Docker image, gateway config
- Homelab skills
- Connect from workstation, verify MCP + SSH
Phase 5 — Secrets (Infisical):
- Stand up Infisical, migrate SOPS, wire services
Phase 6 — Deploy + cutover:
- Docker Compose, Gitea webhook, Caddy re-point
- End-to-end verification
- Stop apps/105, clean up mac-mini
Reuse (logic carried over, rewritten in Go)
| Existing Python | What it becomes in Go |
|---|---|
oikos/decide.py |
internal/policy/classify.go — same scoring logic, reads from DB + patterns |
oikos/signal.py |
internal/scheduler/signal.go — same lifecycle, DB-backed |
oikos/approve.py |
internal/policy/approve.go — same grant lifecycle, DB-backed |
oikos/ledger.py |
internal/db/ — ledger_entries table + sqlc queries |
oikos/policy.py + policy.yaml |
internal/policy/ + seeds/policy.yaml → DB tables |
oikos/drift.py |
internal/scheduler/ — drift detection, writes signals to DB |
oikos/relations.py |
internal/ontology/graph.go — SQL graph traversal |
oikos/report.py |
internal/api/routes/ — report endpoints, reads from DB |
mcp/server.py |
internal/api/mcp.go — MCP adapter on top of DB |
bin/homelab logic |
internal/api/routes/ — same operations, REST interface |
oikos/scheduler.py |
internal/scheduler/ — same probes, goroutines for concurrency |
oikos/approve.py Matrix delivery |
internal/notifier/matrix.go |
| (new) | internal/learning/ — feedback, patterns, skills (no Python equivalent) |
Risks / trade-offs
- Go rewrite — the existing Python code (~4400 lines) is replaced. The logic and design patterns carry over, but it's a full rewrite. Mitigated by the fact that the Python code is well-documented and the Go structure mirrors it.
- Hermes in Docker — agent's world is the container. SSH access is the bridge. Hybrid approach (mounted keys now, actuator gateway later).
- PostgreSQL as SPOF — mitigated by Docker volume persistence + automated
pg_dumpbackups (the scheduler can do this once running). - Learning model cold start — no patterns/skills exist initially. The agent starts cautious (escalates everything), accumulates feedback, and gradually becomes more autonomous as patterns validate. This is by design — trust is earned.
- Infisical bootstrapping — SOPS coexists during transition. Keep SOPS as fallback.
- Multi-agent concurrency —
SELECT ... FOR UPDATE SKIP LOCKEDprevents two actuator passes from acting on the same signal. - Ontology evolution — as the homelab changes, entity types and relationship types need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy.
Audit findings (best practices, security, performance, sanity)
A full audit was performed against the plan. Findings are organized by category and severity. High-severity items must be addressed before implementation; medium items should be addressed during the phase they belong to.
Security
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| S1 | HIGH | SSH private keys mounted into Hermes container — a compromised container has unrestricted SSH to all hosts | Use a dedicated, restricted SSH key with command= in authorized_keys. Time-box the stopgap. Build the actuator gateway (which brokers SSH per-execution) sooner rather than later. |
| S2 | HIGH | MCP interface has no auth — any container on the Docker network has full read access to the control plane | Add a shared secret or mTLS between API and Hermes. Bind MCP to a dedicated Docker network, not the default bridge. Never expose the MCP port via Caddy without auth. |
| S3 | HIGH | Policy DB is mutable via the same API it governs — a compromised API can rewrite its own approval rules (e.g. flip destructive to auto-act) |
Policy mutations require a meta-approval (dual-control). Add an immutable audit log of policy changes. Startup self-check: alert if policy hash differs from a known-good baseline. |
| S4 | HIGH | Learning model poisoning — flapping services or misconfigured probes can inject biased feedback to push patterns past the confidence threshold and unlock auto-act for destructive actions | Require human confirmation before pattern transitions to active. Cap confidence by sample size (require N≥5 executions). Detect anomalous feedback bursts and quarantine. Never let a skill auto-promote to destructive risk class. |
| S5 | HIGH | confirmation_phrase is a weak auth primitive — if reused, stored in plaintext, or transmitted via Matrix, it's replayable |
Make it single-use (one phrase per approval). Store hashed. Transmit only the decision + HMAC, not the phrase. Prefer signed approval tokens. |
| S6 | MEDIUM | REST /exec and /decide endpoints rely on Authentik forward-auth only — if the API port is reachable directly on the mesh, auth disappears |
Enforce OIDC JWT validation in the API middleware too (defense in depth). Bind API port to localhost + Caddy only. |
| S7 | MEDIUM | No TLS between internal services — Postgres connections are plaintext on the Docker network | TLS to Postgres (server cert verification). Use dedicated Docker networks per trust boundary. Document the threat model: is the Docker network trusted? |
| S8 | MEDIUM | Webhook (Gitea → mac-mini deploy) has no auth specified — anyone who can reach the endpoint can trigger arbitrary code execution via go build |
HMAC signature verification on the webhook. Bind listener to localhost (Gitea reaches via mesh). Run deploy script as non-root. Verify commit signatures before git pull. |
| S9 | MEDIUM | Infisical bootstrapping has a chicken-and-egg — Infisical's own master key must come from somewhere | Document the bootstrap root of trust explicitly: where the master key lives (mac-mini keychain), how it's backed up, revocation path. Keep SOPS as fallback until Infisical has a tested restore-from-backup drill. |
| S10 | MEDIUM | Blast radius of a compromised container is wide — Hermes has SSH + MCP + gateway port + data volume; actuator has SSH + writes to policy/learning tables | Principle of least privilege per container: only the actuator should have SSH, not Hermes. Split networks: data (PG), ops (SSH egress), front (Caddy). Use Docker user namespaces and read-only root filesystems. |
Performance
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| P1 | MEDIUM | Recursive CTE blast_radius has no cycle guard — cycles (A→B→A) inflate work exponentially with depth |
Add a visited-set guard (array accumulator in the recursion). Cap max_depth at 3. Add LIMIT on the outer query. Consider a closure table for hot-path queries. |
| P2 | MEDIUM | Concurrent probes with no concurrency cap — unbounded goroutines could exhaust FDs or hammer slow targets | Bounded worker pool (errgroup.SetLimit). Per-target probe timeout. Jitter to avoid thundering herd on shared backends. |
| P3 | MEDIUM | Knowledge graph re-ingestion on every deploy is wasteful if no docs changed | Content-hash each doc (store hash on the entity). Skip ingestion if hash unchanged. Run in a single transaction with deferred FK checks. |
| P4 | MEDIUM | Pattern extraction frequency unspecified — either too tight (scans whole table) or too loose (patterns lag) | Define cadence (hourly). Use a high-watermark on feedback.ts. Add index on feedback(ts). Process only new feedback. |
| P5 | HIGH | Table growth unaddressed — state_snapshots (every entity every 10 min) and executions/feedback grow indefinitely |
Partition state_snapshots and signals by month. Define TTLs (snapshots > 90 days → aggregate, raw → archive). Add a prune_* job in the scheduler. |
| P6 | MEDIUM | WebSocket /events has no backpressure — a stalled client pins a goroutine and accumulates memory |
Bounded channel per subscriber with drop-oldest-on-full. Max subscribers cap. Heartbeat/timeout. Document: best-effort vs. guaranteed delivery. |
| P7 | LOW | Deploy script runs go build on host AND in Docker — double build, host Go toolchain is a deploy dependency |
Standardize on multi-stage Docker build only. Drop host go build from deploy script. Keep host toolchain for local dev only. |
Architecture
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| A1 | HIGH | No testing strategy — 13 workstreams, no mention of unit/integration/property tests | Add a testing workstream: unit tests for classify.go, pattern.go, skill.go; integration tests with testcontainers Postgres; golden-file tests for seed ingest; property test for blast-radius (cycles, depth caps). Coverage gates per package. |
| A2 | HIGH | No observability — no structured logging, metrics, or traces. Can't inspect why the classifier escalated or how long probes took | At minimum: structured JSON logs (slog) with correlation IDs per execution. A /metrics endpoint (even before full Prometheus). debug=true flag for full probe payloads. Reconsider Prometheus deferral — "who watches the watcher" requires metrics. |
| A3 | HIGH | DB backup strategy is a one-liner — the DB is the entire control-plane state (ontology, policy, inventory, learning, ledger) | Define: pg_dump cadence (daily + WAL archiving for PITR), off-host storage (push to hubris or object storage, not the same mac-mini volume), encryption at rest, tested restore procedure (monthly drill), retention (30 daily + 12 monthly). This is Phase 1, not "once running." |
| A4 | MEDIUM | Seed ingest transactional safety — if API crashes mid-ingest, DB could be in partial state | Run migrations + seed ingest as a distinct init container before the API starts. Wrap each seed file in a single transaction. Add a seed_version table to skip already-applied seed versions. |
| A5 | MEDIUM | Config management underspecified — unclear what's env vs. DB vs. file vs. Infisical | Define a config hierarchy: defaults → file → env → Infisical (secrets only). Document each service's required config keys. Avoid putting non-secret config in Infisical. |
| A6 | MEDIUM | apps/105 fallback has no cutover safety — both stacks can't write to the same state without conflict | Define cutover mode: apps/105 is read-only during coexistence. If rollback needed after new OS has been writing to Postgres, document the reconciliation plan. |
| A7 | LOW | Notifier as a separate service is over-engineered initially | Either keep in-process (split later), or document the failure mode: if notifier is down, the API must retry/queue and the operator needs an alternative approval path. |
Data model
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| D1 | HIGH | entities.id as TEXT with manual naming is fragile — renames break the ID; signal IDs require a date-string generator that's race-prone |
Use UUID or SERIAL for PK. Keep name + type as a unique composite for human lookup. Signal IDs: use a DB sequence. |
| D2 | MEDIUM | Ontology flexibility sacrifices type safety — attributes JSONB has no schema enforcement per entity type |
Add a schema JSONB column to entity_types (JSON Schema). Validate attributes against it on insert/update via a trigger or app-layer check. |
| D3 | MEDIUM | Entity type evolution has no migration story — renaming/removing an entity type cascades through many FK references | Add status (active/deprecated) to entity_types. Forbid hard deletes while instances exist. Provide a merge endpoint for renames. |
| D4 | MEDIUM | Concurrent writer safety — counters on patterns (evidence_count, success_count) have read-modify-write races |
Use atomic UPDATE ... SET evidence_count = evidence_count + 1. Add version optimistic-lock column on skills and patterns. |
| D5 | MEDIUM | Migration rollback not addressed — only one .down.sql shown |
Commit to forward-only migrations (compensating migrations for rollbacks) and document it, or write and test every .down.sql. |
| D6 | MEDIUM | Data export for DR is asserted but no endpoint or format defined | Add GET /api/v1/export (returns the three YAML files regenerated from DB). Test round-trip: seed → DB → export → seed → DB yields identical state. |
| D7 | LOW | relationships has no temporal data — can't answer "what depended on authentik last month?" |
Add valid_from/valid_to (nullable = current) for depends-on and hosts edges. Low priority but cheap to add now. |
Operational
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| O1 | HIGH | No rollback strategy — if a deploy breaks the API (bad migration, schema bug), recovery is "revert the commit" but the migration may have already run | Define: pre-deploy DB backup snapshot; migration compatibility policy (new code tolerates old schema for one deploy); tested rollback runbook per phase. |
| O2 | HIGH | "Who watches the watcher" is unresolved — if the OS is down, no alerts fire. No external/orthogonal monitor for the mac-mini or the OS containers | A minimal external watchdog: a cron on apps/105 (or hubris) that curls the API /healthz every 5 min and Matrix-pings the operator directly if it fails. This must be outside the Docker stack. |
| O3 | HIGH | No backup/restore runbook — "pg_dump" is mentioned but no procedure, tested restore, or definition of what "restored" means | Write docs/operations/backup-restore.md: what's backed up (DB, Infisical, Hermes volume, seed YAMLs), where, how often, how to restore each, quarterly restore drill. |
| O4 | HIGH | No disaster recovery plan — if the mac-mini dies (disk, theft, water), what's the RTO/RPO? | Define RTO/RPO targets (homelab: RTO 4h, RPO 24h). Name the off-host backup target. Write the DR runbook: fresh mac-mini → install Docker → clone repo → restore Infisical → restore DB → docker compose up. |
| O5 | MEDIUM | Deploy downtime — docker compose up -d recreates containers; API has a brief gap |
Use docker compose up -d --no-deps for app services. Don't recreate the Postgres container on routine deploys (pin its image). Healthcheck-gated rollout. |
| O6 | MEDIUM | Health checks asserted but not specified | Define per-service: API /healthz (DB ping), scheduler "last successful probe < 15min ago", notifier "last poll < 60s ago", Hermes "gateway responding". Wire into Docker healthcheck + alerting. |
| O7 | MEDIUM | mac-mini single-host failure = total outage — disk failure, macOS update reboot, Docker daemon crash | Automated macOS update deferral/scheduling. Docker restart: always on all services. Monitoring heartbeat from external host. Documented cold-start runbook (what comes up first, in what order). |
Missing
| ID | Severity | Finding | Recommendation |
|---|---|---|---|
| M1 | HIGH | No CI/CD — deploys are git-push → webhook → build. No linting, no tests before deploy, no gated merges | Add CI stage (Gitea Actions): go vet, golangci-lint, go test ./..., docker build (no push). Gate the webhook on green CI. At minimum, deploy script runs go test before docker compose up. |
| M2 | MEDIUM | No rate limiting on the API — a runaway agent loop or misconfigured skill can hammer the API and DB | Per-caller rate limiting (token bucket) on mutating endpoints. /exec needs a per-entity/per-action rate cap to prevent actuator loops. |
| M3 | MEDIUM | No access audit for human activity — the ledger records OS actions, but not operator API actions (entity edits, policy changes, approval decisions) | Add an audit_log table for all mutating REST calls, including operator identity from OIDC. |
| M4 | MEDIUM | No circuit breaker for the actuator — if a target host is unreachable, the actuator keeps attempting SSH and generating failed executions, poisoning the learning model | Circuit breaker per target: after N consecutive failures, back off (exponential) and raise a "target unreachable" signal instead of continuing to execute. |
| M5 | MEDIUM | No secret rotation story — SSH keys, Infisical machine tokens, API shared secrets need rotation policies | Define rotation cadences. Document the rotation procedure for each secret class. Add a "secrets expiring" check to the scheduler. |
| M6 | LOW | No dependency/supply-chain hygiene — Go modules, Docker base images, Infisical image not pinned or scanned | Pin base images by digest. Run govulncheck in CI. Periodically audit go.sum. Use distroless or scratch runtime images. |
| M7 | LOW | No documented SLOs — no quantitative success criteria (probe latency, API p99, deploy time, alert delivery) | Add a small SLO table: probe interval 10min ±1min, API p99 < 200ms, deploy < 5min, alert delivery < 30s. |
Top-priority items to address before implementation
- S1 + S10 — SSH keys in containers / wide blast radius → build the actuator gateway first, not incrementally
- S3 + S4 — Mutable policy DB + learning model poisoning → these compound: a compromised container can rewrite its own rules and inject feedback to unlock auto-act
- A1 + A2 + O2 — No tests, no observability, no external watchdog → can't safely run an autonomous agent without all three
- A3 + O3 + O4 — Backup/restore/DR is a one-liner for the SPOF Postgres
- O1 + M1 — No rollback strategy and no CI gate on deploys
Verification (end to end)
- Ontology:
seeds/ontology.yamlingested —SELECT * FROM entity_typesshows all types across 3 layers; lifecycle definitions match the state machine diagrams. - DB:
docker compose up postgres— all 6 migrations applied (including TimescaleDB extension + hypertables); seed ingest populates entities + relationships frominventory.yaml. - API:
curl http://localhost:8090/api/v1/entities?type=servicereturns the fleet; MCPlist_servicesworks via the same endpoint. - Scheduler: trigger a probe pass — signals in DB, state snapshots written,
metrics in
metric_samples(health, probe_latency_ms, disk_usage_pct). - Actuator: raise a test
service-downsignal → actuator classifies → auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated → event emitted + audit entry written + correlation_id links the full chain. - Learning: after N executions of the same (entity_type, action), a pattern
appears with confidence score; after enough evidence, a skill is created.
pattern_confidencemetric visible in time-series. - Classifier with learning: set
autonomy.auto_act: reversible_low→ next similar signal: classifier checks pattern confidence → auto-acts if high, escalates if low. Kill-switch (auto_act: off) → always escalates. - Hermes: connect from another workstation → agent responds, queries API via
MCP, can SSH to hubris. Agent activity logged to
agent_activitytable. - Secrets: Infisical running, Go services fetch secrets, SOPS files removed.
- Deploy:
git push→ Gitea webhook →docker compose build + up -d→ changes live, seed ingest syncs any YAML changes to DB.deploy.triggered+deploy.completedevents in the event log. - Knowledge: MCP
search_knowledge("caddy")returns docs linked toentity:service:caddy. - Observability:
curl /api/v1/metrics?entity_id=host:hubris&metric=disk_usage_pctreturns time-series with trend.curl /api/v1/audit?entity_id=service:caddyreturns the full audit trail.curl /api/v1/healthreturns fleet summary with trend indicators. MCPget_trend("host:hubris", "disk_usage_pct")returns slope + anomaly detection. - Correlation tracing: follow a
correlation_idfrom signal → classification → execution → SSH command → verification → feedback → pattern update, all linked in the audit_log + events table. - Cutover: stop apps/105, verify production traffic only from Docker OS.
Out of scope (for now)
- Oikos Console web UI (deferred — built later on top of the API)
- Multi-node deployment (designed for, not implemented)
- Vector embeddings / semantic search (structured graph only for now)
- SSH-key-signed approval requests
- Actuator gateway pattern (Phase 2 of hybrid — start with restricted SSH key, build gateway in Phase 3 alongside the actuator)
- Automated skill extraction via LLM (patterns extracted statistically for now; LLM-assisted skill refinement is a future enhancement)
Audit remediation — HIGH priority + architect/developer review
This section addresses all HIGH-severity audit findings and the 18 new findings from the systems architect + senior Go developer review. Each fix references the finding ID.
Schema fixes (CRITICAL/HIGH)
SA1 + SG1 — Cognition entities + broken FKs + hypertable PKs:
The core issue: the BDD models cognition objects (Signal, Execution, Feedback,
Pattern, Skill, Classification) as graph entities with typed edges, but the schema
only makes Signal a dual entity. executions.skill_id REFERENCES entities(id) is a
broken FK (skills live in skills(id), not entities(id)). Additionally, hypertable
PKs (id BIGSERIAL PRIMARY KEY) don't include the time column — TimescaleDB rejects
this.
Resolution — dual entity pattern for all cognition objects:
Every cognition object gets an entities row (type = signal, execution, etc.)
AND a typed table for indexed querying. The typed table's PK references
entities(id). Graph relationships (triggers, produces, contributes-to) live in the
relationships table, making the cognition loop traversable as a graph.
-- Fix: hypertable PKs must include the time column
CREATE TABLE audit_log (
id BIGSERIAL,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
-- ... columns ...
PRIMARY KEY (id, ts) -- was: id BIGSERIAL PRIMARY KEY
);
-- Same fix for events and agent_activity: PRIMARY KEY (id, ts)
-- Fix: executions.skill_id → references skills, not entities
-- (or if using dual-entity pattern, reference entities(id) where type='skill')
CREATE TABLE executions (
id SERIAL PRIMARY KEY,
-- ... existing columns ...
skill_id TEXT, -- entity ID of skill used
skill_version INTEGER, -- SG9: snapshot version at exec time
-- ...
);
-- Fix: skills versioning preserves history (SG9)
-- Change PK to composite so old versions remain queryable
CREATE TABLE skills (
id TEXT NOT NULL, -- 'skill-restart-service'
version INTEGER NOT NULL DEFAULT 1,
name TEXT NOT NULL,
ts TIMESTAMPTZ DEFAULT now(),
procedure TEXT NOT NULL,
applies_to TEXT REFERENCES entity_types(name),
pattern_ids TEXT[],
status TEXT DEFAULT 'drafted',
success_rate REAL,
last_used_at TIMESTAMPTZ,
PRIMARY KEY (id, version) -- was: id TEXT PRIMARY KEY
);
CREATE TABLE skill_versions_audit ( -- SG9: track procedure changes
id SERIAL PRIMARY KEY,
skill_id TEXT NOT NULL,
version INTEGER NOT NULL,
changed_at TIMESTAMPTZ DEFAULT now(),
changed_by TEXT, -- agent or operator entity
diff TEXT, -- diff of procedure field
reason TEXT
);
SA5 — Classification entity must be persisted:
Add a classifications table. Every classifier decision is recorded with the full
reasoning — this is the audit trail for autonomous decisions.
CREATE TABLE classifications (
id SERIAL PRIMARY KEY,
ts TIMESTAMPTZ DEFAULT now(),
signal_entity_id TEXT REFERENCES entities(id),
entity_id TEXT REFERENCES entities(id),
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
route TEXT NOT NULL, -- 'auto-act' or 'escalate'
blast_radius TEXT[],
pattern_confidence REAL,
skill_match TEXT, -- skill entity ID if matched
autonomy_check TEXT, -- 'allowed' or 'blocked: <reason>'
reasoning JSONB NOT NULL, -- full decision explanation
correlation_id TEXT
);
CREATE INDEX idx_class_signal ON classifications(signal_entity_id);
CREATE INDEX idx_class_entity ON classifications(entity_id);
SG2 — CAGG array_agg unsupported:
Drop representative_tags from the 1h continuous aggregate. Query tags from raw
data when needed.
CREATE MATERIALIZED VIEW metric_rollups_1h
WITH (timescaledb.continuous) AS
SELECT
time_bucket('1 hour', ts) AS bucket,
entity_id,
metric,
avg(value) AS avg_value,
min(value) AS min_value,
max(value) AS max_value,
count(*) AS sample_count
-- removed: (array_agg(tags))[1] AS representative_tags
FROM metric_samples
GROUP BY bucket, entity_id, metric;
SG3 — TimescaleDB functions not idempotent:
Use if_not_exists => TRUE and exception guards:
SELECT create_hypertable('metric_samples', 'ts',
chunk_time_interval => INTERVAL '7 days',
if_not_exists => TRUE);
DO $$ BEGIN
PERFORM add_retention_policy('metric_samples', INTERVAL '90 days');
EXCEPTION WHEN OTHERS THEN NULL;
END $$;
Ontology fixes (HIGH)
SA2 — Layer boundary "cognition creates governance" arrow:
Remove the unsupported arrow. The learning engine does NOT write to governance (policy/autonomy) tables. Instead, validated patterns propose autonomy changes as approval-request entities — the operator must accept. This keeps governance authoritative and prevents the learning-poisoning vector (S4).
Also rename the cognition subgraph node from "Governance" to "Approvals" to avoid the layer-name collision.
SA3 — Missing entity types (Person, Agent, IdentityProvider):
Add BDD definitions and entity_type seeds:
Person: matrix_id, oidc_sub
Agent: provider, model, gateway_port
IdentityProvider: issuer, client_id, auth_mode (oidc/forward-auth)
These are first-class governance-layer entities. audit_log.actor_id now resolves
to a real entity.
SA4 — Lifecycle dead-ends and missing transitions:
Updated infrastructure lifecycle (add failed state):
planned → provisioning → active → migrating → deprecated → destroyed
↘ failed ↗ ↘ failed ↗
failed → active (recovery) | failed → deprecated (write-off)
planned → destroyed (cancel)
deprecated → active (un-deprecate if replacement fails)
Updated signal lifecycle (add terminal states):
raised → acknowledged → acting → resolved | failed
acknowledged → resolved (manual resolve without acting)
acknowledged → muted
acting → failed (permanent failure, terminal — needs operator)
failed → acknowledged (operator retries)
Updated execution lifecycle (add cancellation + rollback failure):
approved → expired (TTL ran out)
executing → cancelled (operator abort)
failed → rolled_back | rollback_failed
timed_out → verifying (check if the command actually completed despite timeout)
Updated pattern lifecycle:
hypothesized → validated → active → deprecated
hypothesized → invalidated (disproven, terminal)
active → invalidated (new evidence contradicts)
Updated skill lifecycle:
drafted → tested → active → refined → active (new version)
drafted → deprecated (abandoned)
tested → failed → drafted (back to drawing board)
active → deprecated (superseded or unsafe)
Approval lifecycle (new diagram — was missing):
pending → approved | denied | expired
approved → revoked (operator changes mind before execution)
SA6 — recommended_action data source:
Move recommended_action from signals to classifications. The probe raises a
signal (kind, severity, evidence). The classifier populates the classification with
the recommended action based on signal kind + skill lookup. The actuator reads the
classification, not the signal, for the action to take.
SA8 — Missing entities (Cluster, ComposeStack, ManagedHost):
Add to ontology:
Clusterentity (software domain) —ProxmoxHostmember-ofClusterComposeStackentity (software domain) —DockerContainerpart-ofComposeStack- Add
provider,control_level(full/partial/none) toStandaloneServer
Security remediation (HIGH)
S1 — SSH keys in containers:
Dual approach (operator decision):
- Phase 1-2 (immediate): Restricted SSH key — dedicated key pair with
command="..."andfrom="..."restrictions inauthorized_keyson hubris/ strong. The key can only run specific commands (pct, qm, df, systemctl status), not arbitrary shells. Mounted read-only into the actuator container only (not Hermes). - Phase 3 (actuator build): Full actuator gateway — the API's
/execendpoint brokers all SSH. The actuator holds the keys, executes per-action, logs every command. Hermes never touches SSH.
S2 — MCP auth:
- Shared secret between API and Hermes (HMAC-signed requests)
- MCP bound to a dedicated Docker network (not the default bridge)
- Never exposed via Caddy without auth
S3 — Policy DB mutability:
- Policy mutations (
risk_classes,approval_rules,autonomy_settings) require a meta-approval: the operator must approve the policy change itself (dual-control) - Immutable audit log of all policy changes with before/after hash
- Startup self-check: compute policy hash, alert if differs from last-known-good
S4 — Learning model poisoning:
- Pattern transitions to
activerequire operator confirmation (PATCH /api/v1/ patterns/{id}withstatus=active— policy-gated asconfig_mutation) - Confidence capped by sample size:
confidence = min(raw_confidence, N/5)where N = evidence_count (requires N≥5 for confidence > 0.2) - Anomaly detection: if >10 identical-outcome feedback entries arrive within 1 hour for the same (entity_type, action), quarantine the pattern for review
- Skills can never auto-promote to destructive risk class — always escalate
S5 — confirmation_phrase replacement:
- Replace with single-use signed approval tokens
- Token = HMAC(approval_id + entity_id + action + risk_class + nonce, shared_secret)
- Stored hashed in
approvalstable - Transmitted via Matrix as the approval ID + decision; token verified server-side
SA10 — Gateway + Caddy trust boundary:
- Port 8092 (Hermes gateway): mTLS or token auth. Mesh membership is the network boundary; gateway auth is the application boundary.
- Caddy is an explicit trust root. API validates OIDC JWTs in middleware (not just trusting Caddy headers). Documented: compromising Caddy ≠ compromising the API.
Operational remediation (HIGH)
A3 + O3 + O4 — Backup, restore, DR:
Backup strategy:
- Daily
pg_dump(compressed, custom format) + WAL archiving for PITR - Off-host storage: Proton Drive (cloud object storage, encrypted at rest)
- Push via rclone from the scheduler container (already have rclone LXC in the fleet — reuse credentials)
- Retention: 30 daily + 12 monthly snapshots
- Infisical backup: Infisical has its own backup mechanism; also export secrets to an encrypted SOPS file as a fallback (chicken-and-egg: keep one age key for this purpose)
- Hermes volume: backed up with
pg_dumpof agent_activity + session data
Restore procedure (docs/operations/backup-restore.md):
- Restore Postgres:
pg_restore -d oikos < dump.psql - Verify seed ingest matches (run
GET /api/v1/exportand diff against seed YAML) - Restore Infisical from its backup
docker compose up -d- Monthly restore drill (scheduled, automated, alert if restore fails)
DR plan:
- RTO: 4 hours (fresh machine → Docker → restore → running)
- RPO: 24 hours (last daily backup)
- Cold-start runbook: install Docker → clone repo → restore Infisical →
restore DB →
docker compose up -d→ verify health - Off-host backup target: Proton Drive (encrypted, offsite)
O1 — Rollback strategy:
- Forward-only migrations (no
down.sqlbeyond development). Compensating migrations for production rollbacks. - Pre-deploy DB backup: the deploy script runs
pg_dumpbeforedocker compose up -d - Migration compatibility: new code must tolerate old schema for one deploy window (additive migrations only — new columns nullable, new tables optional)
- Rollback runbook: revert git commit →
pg_restorefrom pre-deploy backup →docker compose up -dwith old image
O2 — External watchdog:
- Cron job on apps/105 (outside the Docker stack):
curl -sf http://mac-mini:8090/healthz || curl -X POST matrix-webhook ... - Runs every 5 minutes
- Alerts operator directly via Matrix if the API is unreachable
- Also checks: Docker daemon running (
docker info), Postgres accepting connections (pg_isready)
M1 — CI/CD:
- Gitea Actions (or simple webhook + script):
go vet ./...golangci-lint rungo test ./... -race -coverdocker build(no push — just verify it builds)
- Webhook deploy gated on green CI
- Deploy script runs
go testas a final safety check beforedocker compose up
Architecture remediation (HIGH)
A1 — Testing strategy:
Add testing workstream with specific tests:
| Package | Test type | What to test |
|---|---|---|
internal/policy/ |
Unit | Classifier scoring: risk × blast × confidence. Table-driven: every (risk_class, blast_radius, confidence) combination. Edge: unknown action, ambiguous entity. |
internal/policy/ |
Unit | Approval lifecycle: token issue, verify, single-use enforcement, TTL expiry. |
internal/learning/ |
Unit | Pattern confidence calculation. Skill versioning. Feedback → pattern extraction. |
internal/ontology/ |
Unit | Lifecycle transition validation: every legal transition succeeds, every illegal one fails. Graph traversal (mock relationships). |
internal/db/ |
Integration | testcontainers Postgres: seed ingest idempotency, blast_radius CTE (with cycles), hypertable insert + query, continuous aggregate refresh. |
internal/api/ |
Integration | testcontainers: REST routes return correct status codes, MCP tools return expected shapes, audit middleware records entries, pagination works. |
internal/actuator/ |
Integration | Mock SSH: execution → verification → feedback → pattern update. Loop-guard prevents retry storms. Circuit breaker trips after N failures. |
migrations/ |
Property | Every migration is forward-only. blast_radius returns correct results on cyclic graphs. Hypertable retention doesn't drop data younger than threshold. |
internal/observability/ |
Unit | Correlation ID propagation through context. Event emitter transactional with state change. Metric recording. |
Coverage gate: ≥80% on internal/policy/ and internal/learning/ (the autonomy-
granting code). ≥60% on everything else.
A2 — Observability: Already addressed (Migration 6, Workstream 14). Update audit status to resolved.
SG4 — Graceful shutdown:
Every cmd/*/main.go implements:
signal.NotifyContext(ctx, syscall.SIGTERM, syscall.SIGINT)- Context propagated to all long-running loops and HTTP server
- Shutdown sequence: stop accepting new work → wait for in-flight (30s deadline) →
for actuator: if execution in-flight, mark
failedwith "shutdown interrupted" + emit feedback → close DB pool docker-compose.yml:stop_grace_period: 30son actuator,stop_signal: SIGTERMon all services
SG5 — Entity-level concurrency:
-- Per-entity advisory lock during execution (prevents concurrent actions
-- on the same target entity, e.g., restart + deploy on the same service)
SELECT pg_advisory_xact_lock(hashtext($1)); -- $1 = target_entity_id
-- ... execute, verify, feedback ...
-- lock released on transaction commit/rollback
SG6 — Domain layer:
Add internal/domain/ package:
internal/domain/
├── entity.go # Entity, EntityType, Relationship domain types
├── signal.go # Signal domain type + lifecycle transition logic
├── execution.go # Execution domain type + state machine
├── classification.go # Classification domain type
├── pattern.go # Pattern domain type + confidence calculation
├── skill.go # Skill domain type + versioning
├── approval.go # Approval domain type + token verification
└── errors.go # Sentinel errors: ErrNotFound, ErrInvalidTransition,
# ErrApprovalRequired, ErrAutonomyBlocked, ErrConflict
DB ↔ domain mapping in internal/db/ (repository pattern). API handlers accept/
return domain types. sqlc models never escape internal/db/.
SG11 — Error handling:
// internal/domain/errors.go
var (
ErrNotFound = errors.New("entity not found")
ErrInvalidTransition = errors.New("invalid lifecycle transition")
ErrApprovalRequired = errors.New("operator approval required")
ErrAutonomyBlocked = errors.New("autonomy policy blocks this action")
ErrConflict = errors.New("concurrent modification conflict")
ErrCircuitOpen = errors.New("circuit breaker open for target")
)
HTTP mapping middleware: ErrNotFound → 404, ErrInvalidTransition → 409,
ErrApprovalRequired → 403, ErrAutonomyBlocked → 403, ErrConflict → 409,
ErrCircuitOpen → 503.
SSH error classification in internal/actuator/execute.go:
- Network unreachable → retryable, circuit breaker
- Auth failure → fatal, alert operator
- Command exit non-zero → execution failed, feedback
- Command timeout → timed_out, feedback
DB retry: serialization failures (SQLSTATE 40001, 40P01) → retry with exponential backoff (max 3 retries).
Go implementation fixes (MEDIUM)
SA9 — TimescaleDB Docker image + migration runner:
- Image:
timescale/timescaledb:2.x-pg16(notpostgres:16) - Migrations run in a one-shot init container (
compose/migrate/Dockerfile) with a dedicated DB user that has DDL but no runtime data privileges - API's DB user gets DML only (least privilege)
golang-migrateGo API with//go:embed migrations/*.up.sql
SG7 — Pattern/skill management endpoints:
PATCH /api/v1/patterns/{id}— state transition (validate/invalidate/deprecate), policy-gated asconfig_mutation, audit-loggedPATCH /api/v1/skills/{id}— state transition + version pin, same gating- This is the operator's manual safety valve for learning-model issues (S4)
SG8 — WebSocket push mechanism:
- In-process event bus (Go channel pub/sub) for events written by the API itself (zero-latency push to WebSocket subscribers)
- Postgres
LISTEN/NOTIFYfor events written by other services (scheduler, actuator) — trigger oneventstable fires NOTIFY after commit - Both feed the WebSocket handler
SG10 — Transactional event emission:
- Event + audit entries written in the same DB transaction as the state change
- If transaction rolls back, events are discarded (never emitted)
LISTEN/NOTIFYfires after commit — subscribers only see committed events
SG13 — Context-aware SSH:
func runSSH(ctx context.Context, client *ssh.Client, cmd string) ([]byte, error) {
session, err := client.NewSession()
if err != nil { return nil, err }
defer session.Close()
type result struct { out []byte; err error }
ch := make(chan result, 1)
go func() { out, err := session.CombinedOutput(cmd); ch <- result{out, err} }()
select {
case r := <-ch: return r.out, r.err
case <-ctx.Done():
session.Close() // unblocks CombinedOutput
client.Close()
return nil, ctx.Err()
}
}
SG14 — Connection pool sizing:
- API: 15 connections, scheduler: 5, actuator: 5, learning: 3 = 28 total
- Postgres
max_connectionsset to 80 - Monitor
db_connections_active, alert if >80% of pool
SG15 — RESTful exec endpoint:
POST /api/v1/executions(wasPOST /api/v1/exec) — creates an execution resource. Handler classifies, checks approval, creates execution inproposedstate.GET /api/v1/executions/{id}— statusPOST /api/v1/executions/{id}/cancel— cancellation
SG16 — Pagination:
- All hypertable-backed endpoints: cursor-based (
?cursor=<ts>&limit=50) entitiesand other small tables: keyset pagination (?after=<id>&limit=50)- Default limit: 50, max: 200
- MCP tools support
limitparameter
SG17 — Go tooling:
sqlc.yamladded to repo layout- Module path:
github.com/dtoro/oikos CGO_ENABLED=0in Dockerfiles,pgx(pure Go),gcr.io/distroless/staticruntime image- Migrations embedded with
//go:embed
SG18 — Health/metrics bypass auth:
/healthzand/metricson a separate Gin router group, no auth, no audit/healthz:SELECT 1against DB/metrics: internal-only (not exposed via Caddy), or token-protected
SA7 — Notifier decoupling:
- API writes
approvalsrow (status=pending) + emitsapproval.requestedevent - Notifier polls pending approvals, sends to Matrix, writes decision directly to
approvalstable (has DB access, not API access) - API polls
approvals.status - No service-to-service calls in either direction — DB is the rendezvous point
- Pending approvals survive Notifier restart
SA9 + SA10 — Threat model documentation:
- Caddy is an explicit trust root (compromise = API compromise, mitigated by JWT validation in API middleware)
- Mesh membership is the network boundary for Hermes gateway
- Docker network is the trust boundary for internal services (mTLS between API and Hermes, TLS to Postgres)
- The actuator is the only container with SSH egress (not Hermes, not the API)
Updated phasing (incorporating remediation)
Phase 0 — Ontology design (no code):
- Finalize entity types (including Person, Agent, IdentityProvider, Cluster, ComposeStack) + relationship types + lifecycles (all with terminal states)
- Write seed manifests
- Review diagrams + lifecycle completeness with operator
Phase 1 — Foundation (Go + DB):
- Go module setup (
github.com/dtoro/oikos), project structure with domain layer - PostgreSQL + TimescaleDB (
timescale/timescaledb:2.x-pg16) - Migrations 1-6 (with fixed hypertable PKs, idempotent TimescaleDB calls, dual entity pattern for cognition objects, classifications table)
- Migration runner as init container (DDL-only DB user)
- Seed ingest pipeline (transactional, init container)
- sqlc + domain layer (repository pattern)
- Structured logging (slog) + error sentinels + HTTP error mapping
- Testing foundation: testcontainers setup, unit test framework, coverage gates
- Backup setup: daily
pg_dump+ rclone push to Proton Drive, WAL archiving - External watchdog: cron on apps/105
Phase 2 — API (Go):
- Gin server with REST routes (resource-oriented, paginated,
/executionsnot/exec) - MCP protocol adapter (shared-secret auth, dedicated network)
- Policy enforcement middleware (OIDC JWT validation, not just Caddy headers)
- Audit middleware (transactional with state changes)
- Event emitter (transactional, in-process bus + LISTEN/NOTIFY)
- Observability routes (metrics, trends, audit, events, health, agent-activity)
- WebSocket (in-process bus + LISTEN/NOTIFY, bounded channels, backpressure)
- Pattern/skill management endpoints (operator override for learning model)
- Domain layer fully fleshed out
- CI: Gitea Actions (go vet, golangci-lint, go test -race -cover, docker build)
Phase 3 — Control loop (Go):
- Scheduler (Observe) — probes, signals, state snapshots, metric recording
- Actuator (Act) — classify, execute, verify, correlation ID propagation
- Restricted SSH key (command= in authorized_keys, actuator-only)
- Entity-level advisory locks (pg_advisory_xact_lock)
- Context-aware SSH (context cancellation, hard timeout)
- Circuit breaker per target host
- Graceful shutdown (SIGTERM, in-flight protection, stop_grace_period: 30s)
- Learning engine — feedback, patterns, skills, learning metrics
- Pattern activation requires operator confirmation (PATCH endpoint)
- Confidence capped by sample size (N≥5)
- Anomaly detection for feedback bursts
- Approval tokens (single-use HMAC, not confirmation phrases)
- Policy meta-approval (dual-control for policy mutations)
- Notifier (DB rendezvous, no service-to-service calls, Matrix impl)
Phase 4 — Agent (Hermes container):
- Hermes Docker image, gateway config (mTLS on port 8092)
- Homelab skills
- Agent activity logging
- Connect from workstation, verify MCP (shared-secret auth) + SSH via actuator
Phase 5 — Secrets (Infisical):
- Stand up Infisical, migrate SOPS, wire services
- Infisical backup + SOPS fallback (one age key kept for DR)
Phase 6 — Deploy + cutover:
- Docker Compose (timescale image, init containers, pool sizing, stop_grace_period)
- Gitea webhook (HMAC auth, non-root deploy user, CI-gated)
- Caddy re-point + JWT validation in API
- End-to-end verification (14 checks)
- Stop apps/105, clean up mac-mini
- Restore drill (monthly, automated)