Files
oikos/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md
dtoro 8850b85325 plans: remediate all HIGH audit items + architect/developer review
Addresses 15 original HIGH audit findings + 18 new findings from
systems architect + senior Go developer review (572 lines added).

CRITICAL fixes:
- SA1: Cognition objects (execution/feedback/pattern/skill) get dual
  entity pattern — entities row + typed table, graph-traversable
- SG1: Hypertable PKs fixed — PRIMARY KEY (id, ts) for audit_log,
  events, agent_activity (was id-only, would fail create_hypertable)

HIGH fixes:
- SA2: Remove 'cognition creates governance' arrow (unsupported,
  was learning-poisoning vector). Patterns propose, operator accepts.
- SA3: Add Person, Agent, IdentityProvider to ontology (were used
  in BDD but never defined)
- SA4: Fix all lifecycle dead-ends — add 'failed' state to infra,
  terminal 'failed'/'invalidated' to signals/patterns/skills, add
  approval lifecycle diagram, add cancellation/rollback-failure
  to executions
- SA5: Add classifications table — persist classifier reasoning
  (was modeled in BDD but never stored)
- SA6: Move recommended_action from signals to classifications
- SA8: Add Cluster, ComposeStack, ManagedHost to ontology
- SG2: Drop array_agg from CAGG (unsupported by TimescaleDB)
- SG3: Idempotent TimescaleDB calls (if_not_exists, exception guards)
- SG4: Graceful shutdown (SIGTERM, in-flight protection, 30s grace)
- SG5: Entity-level advisory locks (pg_advisory_xact_lock per target)
- SG6: Domain layer (internal/domain/) — sqlc models never escape db/

Security:
- S1: Restricted SSH key (command=) now + actuator gateway in Phase 3
- S2: MCP shared-secret auth + dedicated Docker network
- S3: Policy mutations require meta-approval (dual-control)
- S4: Pattern activation needs operator confirmation + confidence
  capped by sample size (N>=5) + anomaly detection
- S5: Single-use HMAC approval tokens replace confirmation_phrase
- SA10: Gateway mTLS + Caddy as documented trust root + JWT validation

Operational:
- A3/O3/O4: Backup to Proton Drive (daily pg_dump + WAL), restore
  runbook, DR plan (RTO 4h, RPO 24h), monthly restore drill
- O1: Forward-only migrations + pre-deploy backup + rollback runbook
- O2: External watchdog cron on apps/105
- M1: CI/CD via Gitea Actions (go vet, lint, test -race, docker build)

Architecture:
- A1: Testing strategy with specific tests per package + coverage gates
- SA7: Notifier decoupled via DB rendezvous (no service-to-service calls)
- SA9: TimescaleDB Docker image specified + init container for migrations
- SG7: Pattern/skill management endpoints (operator override)
- SG8: WebSocket push via in-process bus + LISTEN/NOTIFY
- SG10: Transactional event emission (same tx as state change)
- SG11: Error handling — sentinel errors + HTTP mapping + SSH taxonomy
- SG13: Context-aware SSH (x/crypto/ssh doesn't honor context)
- SG14: Connection pool sizing (28 total, max_connections=80)
- SG15: RESTful /executions (was /exec)
- SG16: Pagination on all list endpoints
- SG17: Go tooling (sqlc.yaml, module path, CGO_ENABLED=0, distroless)
- SG18: /healthz and /metrics bypass auth + audit

Updated phasing incorporates all remediation.
2026-07-06 23:35:50 +02:00

115 KiB
Raw Blame History

Plan: Oikos — Docker-based agentic homelab OS on mac-mini

Status: Planned (2026-07-06, rev 2) — supersedes the launchd-based plan and the Python/Docker rev 1. This revision introduces: Go instead of Python, ontology-first design with systems modeling, DB-native config (inventory/ontology/policy as graph metadata, not YAML files), and a feedback loop where the agent learns from execution.

Vision

Convert this repo into a Docker-based agentic homelab OS written in Go. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors:

  • Operator (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix.
  • Agent — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and learns from every action to improve over time.

All OS services run in Docker containers on mac-mini. The OS is deployed by a git push (Gitea webhook → Docker rebuild). It's designed for mac-mini now, with a path to multi-node later.

Decisions (from operator Q&A, 2026-07-06)

Question Decision
Repo structure One repo, reorganize internally.
Language Go — compiled, type-safe, small containers, goroutines for concurrent probes.
Hermes runtime Runs inside Docker as part of the OS stack (gateway mode).
Agent → homelab access Hybrid — mounted SSH keys now, actuator gateway built incrementally.
Data storage PostgreSQL. Inventory/ontology/policy become DB-native graph metadata.
Config (inventory/ontology/policy) DB-native — YAML files are seed manifests only (bootstrap + DR). The DB is the runtime source of truth. Editable via API, future frontend.
Knowledge/context Structured knowledge graph in Postgres — ontology-defined entity types, typed relationships, linked to operational data.
Operator interface Primary = remote Hermes from any workstation + Matrix. Console/UIs built later for specific tasks.
Host mac-mini for now, designed to scale later.
Deploy Git push → Gitea webhook → Docker rebuild + restart.
bin/homelab CLI Replaced by an API. CLI becomes a thin Go client that calls the OS API.
Secrets Migrate from SOPS+age to Infisical.
Matrix Keep for now, abstract the notification layer for future channels.
MCP server Merged into the unified API — one Go service, REST + MCP interfaces.
Agent type Hermes core + custom homelab skills.
Feedback loop Agent learns from execution — outcomes feed back as patterns and skills that inform future decisions.
Ontology Developed first, before knowledge graph ingestion. Systems modeling: entities, connections, lifecycles.
apps/105 Keep running as fallback until Docker OS is proven.
mac-mini cleanup Start fresh in a new directory, clean up old artifacts later.

Ontology — the systems model

The ontology is the foundational layer of the OS. It defines what exists, how things connect, how they change over time, and how the OS learns about them. It is stored in the database as metadata (entity types, relationship types, lifecycle definitions). YAML seed files bootstrap it on first deploy; after that, the DB is authoritative and editable via API.

Design principles (systems modeling)

  1. Three layers — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's own behavior + learning). Dependencies flow upward: Cognition depends on Governance depends on Infrastructure.
  2. Everything is an entity — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, changes, skills) are first-class entities, not second-class records.
  3. Typed relationships with cardinality — edges carry semantics. hosts is one-to-many; depends-on is many-to-many; documents is one-to-one. The graph is queryable for blast radius, dependency chains, and knowledge lookup.
  4. Lifecycles are state machines — every entity type has a lifecycle. Infrastructure entities move through planned → active → destroyed. Operational entities have their own lifecycles (signals, approvals, patterns, skills). Transitions can require preconditions.
  5. Policy is attached to the ontology — risk classes and approval rules link to entity types and actions. The policy IS part of the model, not a separate file.
  6. Learning is modeled — executions produce outcomes, outcomes accumulate into patterns, patterns refine skills, skills inform future decisions. This is an explicit, queryable part of the graph.

Layer map

graph TB
    cognition -- "observes, acts on, learns about" --> infra
    governance -- "governs access to" --> infra
    governance -- "constrains" --> cognition
    cognition -- "creates + refines" --> governance

    subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"]
        direction LR
        OBS["Observation\nsignal, state-snapshot"]
        DEC["Decision\nclassification, risk-assessment"]
        ACT["Action\nexecution, verification"]
        GOV["Governance\napproval-request, approval-decision"]
        KNOW["Knowledge\ndocument, runbook"]
        LEARN["Learning\npattern, skill, feedback"]
    end

    subgraph governance["Layer 2 — Governance (who controls what)"]
        direction LR
        IDENT["Identity\nperson, agent, identity-provider"]
        SEC["Secrets\nsecret, key, access-grant"]
        POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
    end

    subgraph infra["Layer 1 — Infrastructure (the managed world)"]
        direction LR
        PHYS["Physical\nsite, machine, ups, sensor"]
        COMP["Compute\nmachine, vm, container\n(lxc, docker)"]
        NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
        STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
        SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"]
    end

Block definition diagram (SysML BDD)

Uses Mermaid class diagram syntax following SysML BDD conventions:

  • «abstract» = abstract block (cannot be instantiated)
  • <\|-- = generalization (is-a)
  • *-- = composition (whole-part, lifecycle dependency)
  • o-- = aggregation (whole-part, independent lifecycle)
  • -- = association (typed link)
  • Multiplicity: "1", "0..1", "1..*", "0..*", "*"

Infrastructure layer — compute, storage, network:

classDiagram
    class ComputeEntity {
        <<abstract>>
        +state lifecycle
        +attributes jsonb
    }
    class Machine {
        +cpu_arch
        +ram_gb
    }
    class VirtualMachine {
        +vcpus
        +memory_mb
        +disk_gb
    }
    class Container {
        <<abstract>>
        +runtime
    }
    class LXC {
        +pve_id
        +rootfs
    }
    class DockerContainer {
        +image
        +compose_stack
    }
    class ProxmoxHost {
        +pve_version
        +cluster_member
    }
    class StandaloneServer {
        +hypervisor
    }
    class Workstation {
        +os
        +user
    }
    class Appliance {
        +vendor
        +model
    }
    class Hypervisor {
        +type
        +version
    }

    ComputeEntity <|-- Machine
    ComputeEntity <|-- VirtualMachine
    ComputeEntity <|-- Container
    Machine <|-- ProxmoxHost
    Machine <|-- StandaloneServer
    Machine <|-- Workstation
    Machine <|-- Appliance
    Container <|-- LXC
    Container <|-- DockerContainer

    Machine "1" *-- "0..1" Hypervisor : runs
    Hypervisor "1" o-- "0..*" VirtualMachine : hosts
    Hypervisor "1" o-- "0..*" Container : hosts

    class StoragePool {
        +type lvm, zfs, nfs
        +capacity_gb
    }
    class Volume {
        +name
        +size_gb
    }
    class Mount {
        +mount_point
        +options
    }

    StoragePool "1" *-- "0..*" Volume : contains
    ComputeEntity "1" o-- "0..*" Mount : has
    Mount "0..*" --> "1" Volume : mounts

    class NetworkInterface {
        +mac
        +ip
    }
    class Network {
        <<abstract>>
    }
    class LAN { +subnet}
    class Mesh { +provider}
    class VLAN { +tag}

    ComputeEntity "1" *-- "0..*" NetworkInterface : has
    NetworkInterface "0..*" --> "1" Network : connects-to
    Network <|-- LAN
    Network <|-- Mesh
    Network <|-- VLAN

Software + services layer:

classDiagram
    class Service {
        +port
        +health_url
        +risk_notes
    }
    class Application {
        +version
        +config
    }
    class ConfigRepo {
        +url
        +branch
    }
    class DeployPipeline {
        +trigger
        +target_path
    }
    class IngressRoute {
        +pattern
        +upstream
    }
    class Certificate {
        +issuer
        +expires
    }
    class DNSZone {
        +zone
    }
    class DNSRecord {
        +name
        +record_type
        +value
    }

    ComputeEntity "1" o-- "0..*" Service : provides
    Service "1" *-- "0..*" Application : runs
    Service "0..1" --> "0..1" ConfigRepo : configured-by
    DeployPipeline "0..*" --> "1" Service : deploys-to
    IngressRoute "0..*" --> "1" Service : routes-to
    IngressRoute "0..*" --> "0..1" Certificate : secured-by
    IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by
    Service "0..*" --> "0..*" Service : depends-on
    DNSZone "1" *-- "0..*" DNSRecord : contains
    DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to

Cognition layer — operations + learning:

classDiagram
    class Signal {
        +kind
        +severity
        +state lifecycle
        +evidence
        +recommended_action
    }
    class Execution {
        +status
        +result
        +duration_ms
        +verified
    }
    class Feedback {
        +outcome
        +observation
        +lesson
    }
    class Pattern {
        +confidence
        +evidence_count
        +status
    }
    class Skill {
        +procedure
        +version
        +success_rate
    }
    class Classification {
        +risk
        +route
        +reasoning
    }
    class Approval {
        +status
        +ttl
    }
    class Document {
        +title
        +content
        +source_path
    }
    class Runbook {
        +steps
        +risk_class
        +verification
    }

    ComputeEntity "1" o-- "0..*" Signal : monitored-by
    Signal "0..1" --> "0..*" Execution : triggers
    Classification "1" --> "0..1" Execution : precedes
    Execution "1" *-- "0..1" Feedback : produces
    Feedback "0..*" --> "0..*" Pattern : contributes-to
    Pattern "0..*" --> "0..1" Skill : informs
    Skill "0..1" --> "0..*" Classification : guides
    Execution "0..1" --> "0..1" Approval : requires
    Agent "0..*" --> "0..*" Execution : performs
    Person "0..1" --> "0..*" Approval : decides
    Person "0..1" --> "0..*" Agent : owns

    Entity "1" o-- "0..*" Document : documented-by
    Entity "1" o-- "0..*" Runbook : procedure-for

Design notes — validated assumptions

Can VMs mount storage pools? Yes. In Proxmox, VMs have virtual disks on storage pools (LVM, ZFS, NFS). LXCs have bind mounts and mount points. Both use the same storage pools. The model reflects this: ComputeEntity (abstract) has Mount edges to Volume, regardless of whether the compute entity is a VM, LXC, or machine.

Not every machine is a Proxmox host. The current homelab has: 2 Proxmox hosts, 2 workstations, 1 external VPS, 2 VMs, 19 LXCs. The model uses Machine as the base type with specializations (ProxmoxHost, StandaloneServer, Workstation, Appliance). Only ProxmoxHost runs PVE; StandaloneServer could run KVM/libvirt; Workstation runs desktop OS + optionally Docker. A Hypervisor is software that runs on a Machine and hosts VMs/containers — it's not always present (workstations and appliances may not have one).

Docker containers are first-class compute entities. The OS itself runs in Docker containers, and services like Jellyfin's MariaDB sidecar run in Docker within LXCs. DockerContainer is a specialization of Container with image, compose_stack attributes. This lets the OS model its own infrastructure.

Services run on any compute entity. The homelab has services on LXCs (caddy, gitea), VMs (zimaos, haos), workstations (mac-mini will run the OS), and external hosts (authentik on netbird-vps). The provides relationship is from ComputeEntity (abstract), not from a specific compute type.

Documents and runbooks describe any entity. The old ER diagram had Document → Service only. In practice, docs describe hosts, containers, network infrastructure, storage, and investigations. The model uses Entity (the root abstract type) for documented-by and procedure-for, so any entity can have docs and runbooks.

Infrastructure lifecycle

stateDiagram-v2
    [*] --> planned : operator creates entity
    planned --> provisioning : IP reserved, storage chosen, doc stub
    provisioning --> active : mesh joined, health-check answering, doc complete
    active --> migrating : preflight + backup verified
    migrating --> active : post-verify, caddy checked, mounts checked
    active --> deprecated : replacement live or role retired
    deprecated --> destroyed : backups verified, secrets revoked, ingress removed
    destroyed --> [*] : archaeology entry recorded

    note right of deprecated
        Complete only when
        zero inbound edges remain
        (no depends-on, no routes-to)
    end note

Signal lifecycle

stateDiagram-v2
    [*] --> raised : scheduler probe or agent finding
    raised --> acknowledged : agent or operator sees it
    raised --> muted : operator suppresses (TTL)
    acknowledged --> acting : actuator starts execution
    acting --> resolved : action succeeded, verification passed
    acting --> raised : action failed, re-escalated
    raised --> resolved : condition cleared (auto-resolve)
    muted --> raised : TTL expired
    resolved --> [*]

Execution + learning lifecycle

stateDiagram-v2
    [*] --> proposed : signal + recommended_action
    proposed --> approved : operator approves (if gated)
    proposed --> auto_approved : risk class allows auto-act
    approved --> executing : actuator runs
    auto_approved --> executing : actuator runs
    executing --> verified : verification command passed
    executing --> failed : execution or verification failed
    executing --> timed_out : exceeded duration limit
    failed --> rolled_back : rollback procedure executed
    verified --> [*] : feedback recorded, pattern updated
    failed --> [*] : feedback recorded, pattern updated
    rolled_back --> [*] : feedback recorded, pattern updated
    timed_out --> [*] : feedback recorded, pattern updated

Feedback / learning model — the cognition loop

flowchart TB
    subgraph observe["Observe"]
        SIG["Signal raised\n(service down, disk full, drift)"]
    end

    subgraph decide["Decide"]
        CLASS["Classification\nrisk × blast radius × confidence"]
        SKILL_LOOKUP["Skill lookup\nbest-known procedure for\nthis entity type + action"]
        CLASS --> SKILL_LOOKUP
    end

    subgraph act["Act"]
        EXEC["Execution\nfollow skill procedure\n(or escalate if no skill)"]
        VERIFY["Verification\ncheck if action succeeded"]
        EXEC --> VERIFY
    end

    subgraph learn["Learn (feedback loop)"]
        OUTCOME["Outcome evaluation\nsuccess / failure / partial / unexpected"]
        FEEDBACK["Feedback record\nwhat happened vs expected\nextractable lesson"]
        PATTERN["Pattern extraction\naccumulate feedback on\nsimilar entity+action pairs"]
        SKILL_REFINE["Skill refinement\nupdate or create skill\nbased on validated patterns"]

        OUTCOME --> FEEDBACK
        FEEDBACK --> PATTERN
        PATTERN --> SKILL_REFINE
    end

    SIG --> CLASS
    SKILL_LOOKUP --> EXEC
    VERIFY --> OUTCOME
    SKILL_REFINE -. "informs next decision" .-> SKILL_LOOKUP

    subgraph escalation["Escalation path"]
        APPROVAL["Approval request\n→ Matrix ✅/❌"]
    end

    CLASS -- "needs approval" --> APPROVAL
    APPROVAL -- "approved" --> EXEC
    APPROVAL -- "denied" --> RESOLVE["Resolve signal\nnote: denied"]

Policy model — how risk attaches to entities

flowchart LR
    subgraph ontology["Ontology (DB metadata)"]
        ET["entity_types\nhost, service, lxc, vm, ..."]
        RT["relationship_types\nhosts, provides, depends-on, ..."]
        LD["lifecycle_defs\nstates + transitions"]
    end

    subgraph policy["Policy (DB records)"]
        RC["risk_classes\nread_only, reversible_low,\nconfig_mutation, destructive"]
        RULES["approval_rules\nentity_type + action → risk_class\n+ approval_required + autonomy_level"]
        AUTO["autonomy_settings\nglobal: auto_act on/off\nper-entity: never_auto_act"]
    end

    subgraph instances["Instances (DB data)"]
        ENT["entities\nhost:hubris, service:caddy, ..."]
        REL["relationships\nhubris hosts lxc:apps"]
    end

    ET --> RULES
    RC --> RULES
    RULES --> ENT
    AUTO --> ENT
    ET --> ENT
    RT --> REL
    LD --> ENT

Target architecture

Container stack on mac-mini

graph TB
    subgraph mac-mini["mac-mini — Docker host, always-on"]
        subgraph services["Docker Compose — Go binaries"]
            PG["PostgreSQL 16 + TimescaleDB\n• entity/relationship store\n• signals • ledger\n• knowledge graph\n• patterns + skills\n• policy + ontology metadata\n• metric_samples (hypertable)\n• audit_log (hypertable)\n• events (hypertable)\n• agent_activity (hypertable)"]
            INF["Infisical\n(secrets manager)"]
            HERMES["Hermes Agent — gateway mode\n+ homelab skills\n+ MCP client → API\n+ SSH keys (mounted)"]
            API["Oikos API — Go (Gin)\n\nMCP: get_host, list_services,\nsearch_knowledge, get_relations,\nquery_metrics, get_trend,\nget_audit_trail, get_health_summary\nREST: /hosts, /services, /signals,\n/exec, /approve, /deploy, /events,\n/metrics, /audit, /health\n\nPolicy enforcement + risk\nclassification + ledger\nAudit middleware + event emitter"]
            SCHED["Scheduler (Observe) — Go\n+ Actuator (Act) — Go\n• 10-min probes → DB + metrics\n• Classify → auto-act or escalate\n• Execution → feedback → patterns\n• Correlation ID propagation\n• SSH to hubris/strong"]
            NOTIFIER["Notifier — Go\n• Matrix (current)\n• Future: webhook, email"]
        end
        DEPLOY["Gitea webhook →\ndocker compose build + up -d"]
    end

    HERMES -- MCP --> API
    API --> PG
    SCHED --> PG
    SCHED -- "metrics + events + audit" --> PG
    API -- "audit + events" --> PG
    HERMES -- "agent_activity" --> PG
    SCHED -- SSH --> HUBRIS
    SCHED -- SSH --> STRONG
    API -- "escalate" --> NOTIFIER
    NOTIFIER -- "alerts + approvals" --> MATRIX
    DEPLOY -- "rebuild" --> services

    subgraph external["External"]
        CADDY["Caddy (LXC 121)\n→ mac-mini mesh :8090"]
        APPS["apps/105 (fallback)"]
        HUBRIS["hubris (PVE)"]
        STRONG["strong (PVE)"]
        GITEA["Gitea (LXC 104)"]
        MATRIX["Matrix (LXC 118)"]
        WS["Any workstation\nHermes remote → gateway"]
    end

    CADDY -- reverse_proxy --> API
    GITEA -- webhook --> DEPLOY
    WS -- "Hermes gateway" --> HERMES

Observability — data capture and query flow

flowchart TB
    subgraph sources["Data sources"]
        SCHED_PROBE["Scheduler probes\nhealth, disk, latency, drift"]
        API_CALLS["API calls\nrequest/response metrics"]
        AGENT_CALLS["Agent activity\nMCP tools, SSH, reasoning"]
        EXEC["Executions\nSSH commands, verification"]
        LEARN["Learning engine\npatterns, skills, confidence"]
        DEPLOY_EVT["Deploy events\nwebhook, build, restart"]
    end

    subgraph capture["Capture layer (Go)"]
        METRICS_W["metrics.go\n→ metric_samples"]
        AUDIT_W["audit.go\n→ audit_log"]
        EVENTS_W["events.go\n→ events table"]
        AGENT_W["agent.go\n→ agent_activity"]
    end

    subgraph store["PostgreSQL + TimescaleDB"]
        MS["metric_samples\n(hypertable, 90d raw,\n1y rollups via CAGGs)"]
        AL["audit_log\n(hypertable, 1y retention)"]
        EV["events\n(hypertable, 90d retention)"]
        AA["agent_activity\n(hypertable, 90d retention)"]
    end

    subgraph query["Query layer (API)"]
        REST["REST endpoints\n/metrics, /audit,\n/events, /health,\n/trends, /agent-activity"]
        MCP["MCP tools\nquery_metrics, get_trend,\nget_audit_trail, get_event_timeline,\nget_agent_activity, get_health_summary"]
        WS["WebSocket\n/api/v1/events\n(live stream)"]
    end

    subgraph consumers["Consumers"]
        AGENT["Agent\nqueries trends, audits\nits own history"]
        FUTURE["Future: Grafana,\ndashboards, notebooks,\nSIEM, compliance tools"]
    end

    SCHED_PROBE --> METRICS_W
    API_CALLS --> METRICS_W
    API_CALLS --> AUDIT_W
    AGENT_CALLS --> AGENT_W
    EXEC --> EVENTS_W
    EXEC --> AUDIT_W
    LEARN --> METRICS_W
    LEARN --> EVENTS_W
    DEPLOY_EVT --> EVENTS_W

    METRICS_W --> MS
    AUDIT_W --> AL
    EVENTS_W --> EV
    AGENT_W --> AA

    MS --> REST
    AL --> REST
    EV --> REST
    EV --> WS
    AA --> REST
    MS --> MCP
    AL --> MCP
    EV --> MCP
    AA --> MCP

    REST --> AGENT
    MCP --> AGENT
    REST --> FUTURE
    WS --> FUTURE

OODA loop — with learning feedback

flowchart LR
    OBSERVE["Observe\nScheduler probes:\n• HTTP health\n• disk usage\n• drift detection"] --> ORIENT["Orient\nRelations graph walk:\n• blast radius\n• lifecycle state\n• runbook match"]
    ORIENT --> DECIDE["Decide\nRisk classifier +\nskill lookup:\nrisk × blast × confidence"]
    DECIDE -- "auto-act" --> ACT["Act\nExecute via SSH\n→ verify → feedback"]
    DECIDE -- "escalate" --> APPROVE["Approval\n→ Matrix ✅/❌"]
    APPROVE -- "approved" --> ACT
    ACT --> LEARN["Learn\nOutcome → feedback\n→ pattern → skill"]
    LEARN -. "improves confidence" .-> DECIDE
    LEARN --> OBSERVE

Deploy flow

flowchart LR
    DEV["Operator\nedits repo"] --> PUSH["git push"] --> GITEA["Gitea\n(LXC 104)"]
    GITEA -- "webhook" --> MACMINI["mac-mini\ndeploy script"]
    MACMINI -- "git pull" --> REPO["repo clone"]
    MACMINI -- "go build +\ndocker compose up -d" --> STACK["OS containers\nrebuilt + restarted"]
    STACK -- "seed ingest" --> DB["PostgreSQL\nontology + inventory + policy\nsynced from YAML seeds"]

DB-native configuration

The three YAML files — inventory.yaml, ontology.yaml, policy.yaml — become seed manifests. They bootstrap the DB on first deploy. After that, the DB is the runtime source of truth, editable via the API. A future frontend can edit all three directly.

How it works

flowchart TB
    subgraph seeds["Seed manifests (git-tracked, YAML)"]
        ONTO_YAML["seeds/ontology.yaml\nentity types, relationship types,\nlifecycle definitions"]
        INV_YAML["seeds/inventory.yaml\nentity instances (hosts, services,\nnetworks, storage)"]
        POL_YAML["seeds/policy.yaml\nrisk classes, approval rules,\nautonomy settings"]
    end

    subgraph db["PostgreSQL (runtime source of truth)"]
        META["entity_types table\nrelationship_types table\nlifecycle_defs table"]
        INST["entities table\nrelationships table"]
        POLDB["policies table\nrisk_classes table\nautonomy_settings table"]
    end

    INGEST["Seed ingest (on deploy)\nidempotent upsert"]
    ONTO_YAML --> INGEST --> META
    INV_YAML --> INGEST --> INST
    POL_YAML --> INGEST --> POLDB

    API_EDIT["API edits\n(POST/PUT/PATCH)"]
    API_EDIT --> META
    API_EDIT --> INST
    API_EDIT --> POLDB

    EXPORT["Export to YAML\n(for DR / version control)"]
    META --> EXPORT
    INST --> EXPORT
    POLDB --> EXPORT

Why DB-native

  • Querying — the agent can ask "what services depend on authentik?" as a graph query, not a YAML parse. Blast-radius walks are SQL, not file reads.
  • Mutation — adding a service, updating a lifecycle state, changing a policy rule are DB transactions with audit trail, not file edits + git commits.
  • Consistency — the ontology, inventory, and policy are always in sync (same DB, same transaction). No drift between what the YAML says and what the runtime sees.
  • Future frontend — a UI can edit entities, relationships, and policies directly via the API. No need to generate/edit YAML files.
  • Version control — seed YAML files are still git-tracked for bootstrap and DR. The API can export the current DB state back to YAML for commit.

Repo layout (Go project)

/                              # repo root
├── docker-compose.yml         # the OS stack definition
├── Makefile                   # build, test, deploy targets
├── go.mod                     # Go module definition
├── go.sum
├── cmd/                       # binary entrypoints (one per service)
│   ├── api/                   # Oikos API server
│   │   └── main.go
│   ├── scheduler/             # Observe + Act loop
│   │   └── main.go
│   └── notifier/              # Notification service
│       └── main.go
├── internal/                  # private packages (not importable)
│   ├── db/                    # database layer
│   │   ├── queries/           # sqlc SQL queries
│   │   ├── models.go          # generated Go types
│   │   └── db.go              # connection pool, migrations
│   ├── ontology/              # ontology types + meta-schema
│   │   ├── types.go           # EntityType, RelationshipType, LifecycleDef
│   │   ├── graph.go           # graph traversal (blast radius, dependencies)
│   │   └── ingest.go          # YAML seed → DB ingest
│   ├── api/                   # HTTP + MCP server
│   │   ├── server.go          # Gin app setup
│   │   ├── routes/            # REST handlers
│   │   │   ├── hosts.go
│   │   │   ├── services.go
│   │   │   ├── signals.go
│   │   │   ├── approvals.go
│   │   │   ├── exec.go
│   │   │   └── knowledge.go
│   │   └── mcp.go             # MCP protocol adapter (JSON-RPC over SSE)
│   ├── policy/                # risk classification + approval
│   │   ├── classify.go        # risk × blast × confidence
│   │   ├── approve.go         # approval request + grant lifecycle
│   │   └── autonomy.go        # kill-switch, never-auto-act list
│   ├── scheduler/             # Observe stage
│   │   ├── probe.go           # HTTP health, disk, drift
│   │   └── signal.go          # raise/resolve signals in DB
│   ├── actuator/              # Act stage
│   │   ├── act.go             # read signals, classify, execute or escalate
│   │   ├── execute.go         # SSH execution + verification
│   │   └── guard.go           # loop-guard, retry caps
│   ├── learning/              # feedback loop (the learning model)
│   │   ├── feedback.go        # record outcome + lesson from execution
│   │   ├── pattern.go         # extract/validate patterns from feedback
│   │   └── skill.go           # create/refine skills from patterns
│   ├── notifier/              # notification abstraction
│   │   ├── notifier.go        # interface
│   │   └── matrix.go          # Matrix implementation
│   ├── observability/         # logging, metrics, audit, events (NEW)
│   │   ├── logging.go         # slog structured logging setup
│   │   ├── metrics.go         # metric recording (writes to metric_samples)
│   │   ├── audit.go           # audit middleware (writes to audit_log)
│   │   ├── events.go          # event emitter (writes to events table)
│   │   ├── agent.go           # agent activity recording
│   │   └── correlation.go     # correlation ID propagation (context-based)
│   └── config/                # config loading (env, files)
│       └── config.go
├── migrations/                # SQL migrations (golang-migrate format)
│   ├── 001_ontology.up.sql    # meta-schema (entity_types, relationship_types, ...)
│   ├── 001_ontology.down.sql
│   ├── 002_instances.up.sql   # entities, relationships
│   ├── 003_operations.up.sql  # signals, approvals, executions, feedback
│   ├── 004_learning.up.sql    # patterns, skills
│   ├── 005_policy.up.sql      # policies, risk_classes, autonomy
│   └── 006_observability.up.sql # metrics, audit_log, events, agent_activity
├── seeds/                     # YAML seed manifests (bootstrap + DR)
│   ├── ontology.yaml          # entity types, relationship types, lifecycles
│   ├── inventory.yaml         # entity instances (hosts, services, etc.)
│   └── policy.yaml            # risk classes, approval rules, autonomy
├── docs/                      # narrative docs (ingested into knowledge graph)
│   ├── containers/
│   ├── hosts/
│   ├── infrastructure/
│   └── investigations/
├── hermes/                    # Hermes agent config + skills
│   ├── config.yaml
│   ├── SOUL.md
│   └── skills/
│       └── homelab-ops/
│           └── SKILL.md
├── compose/                   # Docker build contexts
│   ├── api/Dockerfile
│   ├── scheduler/Dockerfile
│   ├── hermes/Dockerfile
│   └── postgres/init.sql
└── scripts/                   # utility scripts
    ├── migrate-sops.sh        # one-time SOPS → Infisical migration
    └── import-legacy.sh       # import existing signals/ledger JSONL

Database schema

The schema is the ontology made concrete. Five migration groups, each adding a layer.

Migration 1: Ontology meta-schema

-- The meta-graph: defines what entity types and relationship types can exist.
-- This IS the ontology, stored in the DB, editable via API.

CREATE TABLE entity_types (
    name          TEXT PRIMARY KEY,         -- 'host', 'service', 'signal', 'pattern'
    domain        TEXT NOT NULL,            -- 'physical', 'compute', 'network', ...
    layer         TEXT NOT NULL,            -- 'infrastructure', 'governance', 'cognition'
    description   TEXT,
    lifecycle_id  TEXT,                     -- FK to lifecycle_defs (nullable = no lifecycle)
    attribute_schema JSONB,                -- JSON Schema for validating entity.attributes
    status        TEXT NOT NULL DEFAULT 'active',  -- 'active', 'deprecated' (no hard delete while instances exist)
    created_at    TIMESTAMPTZ DEFAULT now(),
    updated_at    TIMESTAMPTZ DEFAULT now()
);

CREATE TABLE relationship_types (
    name          TEXT PRIMARY KEY,         -- 'hosts', 'provides', 'depends-on'
    inverse       TEXT,                     -- 'runs-on', 'provided-by'
    source_type   TEXT REFERENCES entity_types(name),
    target_type   TEXT REFERENCES entity_types(name),
    cardinality   TEXT NOT NULL,            -- 'one-to-one', 'one-to-many', 'many-to-many'
    description   TEXT,
    created_at    TIMESTAMPTZ DEFAULT now()
);

CREATE TABLE lifecycle_defs (
    id            TEXT PRIMARY KEY,         -- 'infrastructure', 'signal', 'execution', ...
    states        TEXT[] NOT NULL,           -- ordered states
    default_state TEXT NOT NULL,
    transitions   JSONB NOT NULL,            -- {"from": {"to": {"requires": [...]}}}
    created_at    TIMESTAMPTZ DEFAULT now()
);

Migration 2: Entity instances (the inventory graph)

-- Entity instances — the actual hosts, services, signals, patterns, etc.
-- This replaces inventory.yaml as the runtime source of truth.

CREATE TABLE entities (
    id            TEXT PRIMARY KEY,         -- 'host:hubris', 'service:caddy', 'sig:2026-07-06-0001'
    type          TEXT NOT NULL REFERENCES entity_types(name),
    name          TEXT NOT NULL,             -- 'hubris', 'caddy', 'disk-threshold'
    state         TEXT,                      -- lifecycle state (e.g. 'active', 'raised')
    attributes    JSONB NOT NULL DEFAULT '{}',  -- type-specific data (IP, mesh addr, port, ...)
    parent_id     TEXT REFERENCES entities(id), -- for hierarchical entities (LXC on host)
    created_at    TIMESTAMPTZ DEFAULT now(),
    updated_at    TIMESTAMPTZ DEFAULT now()
);

CREATE INDEX idx_entities_type ON entities(type);
CREATE INDEX idx_entities_state ON entities(state);
CREATE INDEX idx_entities_attributes ON entities USING GIN(attributes);

-- Relationship instances — the typed edges of the graph
CREATE TABLE relationships (
    source_id     TEXT NOT NULL REFERENCES entities(id),
    target_id     TEXT NOT NULL REFERENCES entities(id),
    type          TEXT NOT NULL REFERENCES relationship_types(name),
    attributes    JSONB,
    created_at    TIMESTAMPTZ DEFAULT now(),
    PRIMARY KEY (source_id, target_id, type)
);

CREATE INDEX idx_rel_source ON relationships(source_id);
CREATE INDEX idx_rel_target ON relationships(target_id);
CREATE INDEX idx_rel_type ON relationships(type);

-- Recursive graph traversal function (blast radius, dependency chains)
CREATE OR REPLACE FUNCTION blast_radius(start_id TEXT, max_depth INT DEFAULT 3)
RETURNS TABLE(entity_id TEXT, depth INT) AS $$
    WITH RECURSIVE walk AS (
        SELECT start_id::TEXT AS entity_id, 0::INT AS depth
        UNION
        SELECT r.target_id::TEXT, w.depth + 1
        FROM relationships r
        JOIN walk w ON r.source_id = w.entity_id
        WHERE w.depth < max_depth
    )
    SELECT DISTINCT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
$$ LANGUAGE sql STABLE;

Migration 3: Operations (signals, approvals, ledger, state)

-- Signals — now entities in the graph, with a dedicated table for indexed querying
-- (the entity row is the canonical record; this table is a fast lookup)
CREATE TABLE signals (
    entity_id         TEXT PRIMARY KEY REFERENCES entities(id),
    kind              TEXT NOT NULL,
    severity          TEXT NOT NULL,            -- info, warning, critical
    target_entity_id  TEXT REFERENCES entities(id),  -- the infrastructure entity this is about
    evidence          TEXT,
    likely_cause      TEXT,
    recommended_action JSONB,
    verification      TEXT,
    state             TEXT NOT NULL DEFAULT 'raised',
    mute_until        TIMESTAMPTZ,
    created_at        TIMESTAMPTZ DEFAULT now(),
    updated_at        TIMESTAMPTZ DEFAULT now()
);

CREATE INDEX idx_signals_state ON signals(state);
CREATE INDEX idx_signals_target ON signals(target_entity_id);
CREATE INDEX idx_signals_severity ON signals(severity);

-- Approvals
CREATE TABLE approvals (
    id                TEXT PRIMARY KEY,
    ts                TIMESTAMPTZ DEFAULT now(),
    entity_id         TEXT REFERENCES entities(id),       -- entity to act on
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    status            TEXT NOT NULL DEFAULT 'pending',    -- pending, approved, denied, expired
    ttl               INTERVAL NOT NULL DEFAULT '1 hour',
    decided_at        TIMESTAMPTZ,
    decided_by        TEXT REFERENCES entities(id),       -- person entity
    confirmation_phrase TEXT
);

-- Change ledger (high-level record, links to execution for detail)
CREATE TABLE ledger_entries (
    id                SERIAL PRIMARY KEY,
    ts                TIMESTAMPTZ DEFAULT now(),
    entity_id         TEXT REFERENCES entities(id),
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    result            TEXT,                            -- ok, failed, escalated
    approval_id       TEXT REFERENCES approvals(id),
    execution_id      INTEGER,                         -- FK to executions (migration 4)
    agent_id          TEXT REFERENCES entities(id),
    notes             TEXT
);

-- State snapshots (replaces oikos/state.json)
CREATE TABLE state_snapshots (
    id                SERIAL PRIMARY KEY,
    ts                TIMESTAMPTZ DEFAULT now(),
    entity_id         TEXT REFERENCES entities(id),
    health            TEXT,                            -- healthy, degraded, down, unknown
    data              JSONB
);

Migration 4: Learning model (executions, feedback, patterns, skills)

-- Executions — detailed record of each action the OS performs
CREATE TABLE executions (
    id                SERIAL PRIMARY KEY,
    ts                TIMESTAMPTZ DEFAULT now(),
    signal_entity_id  TEXT REFERENCES entities(id),    -- signal that triggered this
    target_entity_id  TEXT REFERENCES entities(id),    -- entity acted upon
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    approval_id       TEXT REFERENCES approvals(id),
    agent_id          TEXT REFERENCES entities(id),    -- who/what executed
    skill_id          TEXT REFERENCES entities(id),    -- skill used (if any)
    status            TEXT NOT NULL DEFAULT 'queued',  -- queued, running, completed, failed, timed-out
    result            JSONB,                           -- detailed result data
    duration_ms       INTEGER,
    verified          BOOLEAN DEFAULT false,
    started_at        TIMESTAMPTZ,
    completed_at      TIMESTAMPTZ
);

CREATE INDEX idx_exec_target ON executions(target_entity_id);
CREATE INDEX idx_exec_status ON executions(status);
CREATE INDEX idx_exec_action ON executions(action);

-- Feedback — what was learned from an execution
CREATE TABLE feedback (
    id                SERIAL PRIMARY KEY,
    execution_id      INTEGER REFERENCES executions(id),
    ts                TIMESTAMPTZ DEFAULT now(),
    outcome           TEXT NOT NULL,                   -- success, failure, partial, unexpected
    observation       TEXT,                            -- what happened vs what was expected
    lesson            TEXT,                            -- extractable lesson
    unexpected_side_effects TEXT[],
    tags              TEXT[]
);

CREATE INDEX idx_feedback_execution ON feedback(execution_id);
CREATE INDEX idx_feedback_outcome ON feedback(outcome);

-- Patterns — generalized rules extracted from accumulated feedback
CREATE TABLE patterns (
    id                TEXT PRIMARY KEY,                -- 'pat-2026-07-06-001'
    ts                TIMESTAMPTZ DEFAULT now(),
    entity_type       TEXT REFERENCES entity_types(name),  -- applies to this type
    action            TEXT NOT NULL,                   -- 'restart', 'deploy', etc.
    pattern           TEXT NOT NULL,                   -- 'service X recovers within 30s after restart'
    confidence        REAL DEFAULT 0.5,                -- 0.0 to 1.0
    evidence_count    INTEGER DEFAULT 1,               -- how many executions support this
    success_count     INTEGER DEFAULT 0,
    failure_count     INTEGER DEFAULT 0,
    status            TEXT DEFAULT 'hypothesized',     -- hypothesized, validated, active, deprecated
    last_validated_at TIMESTAMPTZ
);

CREATE INDEX idx_patterns_type_action ON patterns(entity_type, action);
CREATE INDEX idx_patterns_status ON patterns(status);

-- Skills — codified procedures refined through feedback
CREATE TABLE skills (
    id                TEXT PRIMARY KEY,                -- 'skill-restart-service', 'skill-deploy-lxc'
    name              TEXT NOT NULL,
    ts                TIMESTAMPTZ DEFAULT now(),
    procedure         TEXT NOT NULL,                   -- the codified steps (markdown or structured)
    applies_to        TEXT REFERENCES entity_types(name),
    pattern_ids       TEXT[],                          -- patterns that inform this skill
    status            TEXT DEFAULT 'drafted',          -- drafted, tested, active, refined, deprecated
    version           INTEGER DEFAULT 1,
    success_rate      REAL,                            -- rolling success rate
    last_used_at      TIMESTAMPTZ
);

CREATE INDEX idx_skills_type ON skills(applies_to);
CREATE INDEX idx_skills_status ON skills(status);

Migration 5: Policy (DB-native risk + approval rules)

-- Risk classes — the four-level safety model
CREATE TABLE risk_classes (
    name              TEXT PRIMARY KEY,                -- 'read_only', 'reversible_low', etc.
    description       TEXT,
    approval_required TEXT NOT NULL DEFAULT 'none',    -- none, operator, operator_confirmed
    ledger            BOOLEAN DEFAULT false,
    autonomy_allowed  BOOLEAN DEFAULT false            -- can agent auto-act at this risk level?
);

-- Approval rules — entity_type + action → risk_class + requirements
CREATE TABLE approval_rules (
    id                SERIAL PRIMARY KEY,
    entity_type       TEXT REFERENCES entity_types(name),  -- applies to this entity type
    action            TEXT NOT NULL,                        -- 'restart', 'deploy', 'destroy'
    risk_class        TEXT NOT NULL REFERENCES risk_classes(name),
    autonomy_level    TEXT NOT NULL DEFAULT 'auto',        -- 'auto', 'escalate', 'never'
    scope_entity      TEXT REFERENCES entities(id),        -- optional: specific entity only
    created_at        TIMESTAMPTZ DEFAULT now(),
    updated_at        TIMESTAMPTZ DEFAULT now(),
    UNIQUE(entity_type, action)
);

-- Autonomy settings — global kill-switch + per-entity overrides
CREATE TABLE autonomy_settings (
    key               TEXT PRIMARY KEY,                -- 'global.auto_act', 'never_auto_act.caddy'
    value             TEXT NOT NULL,                   -- 'off', 'reversible_low', 'true', 'false'
    updated_at        TIMESTAMPTZ DEFAULT now()
);

Migration 6: Observability (metrics, audit log, event log)

Uses TimescaleDB (PostgreSQL extension) for time-series data. Hypertables auto-partition by time, continuous aggregates provide rollups, and retention policies auto-drop old data. No separate database needed — everything stays in Postgres.

-- Enable TimescaleDB
CREATE EXTENSION IF NOT EXISTS timescaledb;

-- ─── Time-series metrics ──────────────────────────────────────────────
-- Generic metric store. Every probe, health check, and system measurement
-- writes here. Designed for high insert volume, time-range queries, and
-- continuous-aggregate rollups.

CREATE TABLE metric_samples (
    ts          TIMESTAMPTZ NOT NULL,
    entity_id   TEXT NOT NULL,                 -- which entity this metric is about
    metric      TEXT NOT NULL,                 -- 'health', 'disk_usage_pct', 'probe_latency_ms',
                                               -- 'api_p99_ms', 'goroutines', 'db_connections',
                                               -- 'pattern_confidence', 'skill_success_rate', ...
    value       DOUBLE PRECISION NOT NULL,
    tags        JSONB DEFAULT '{}'::JSONB      -- arbitrary key-value labels:
                                               -- {probe: "http", target: "192.168.8.121"},
                                               -- {host: "hubris", mount: "/mnt/library"}, ...
);

-- Hypertable: partition by time, 1-week chunks
SELECT create_hypertable('metric_samples', 'ts', chunk_time_interval => INTERVAL '7 days');

-- Indexes for common query patterns
CREATE INDEX idx_metrics_entity_ts ON metric_samples(entity_id, ts DESC);
CREATE INDEX idx_metrics_metric_ts ON metric_samples(metric, ts DESC);
CREATE INDEX idx_metrics_tags     ON metric_samples USING GIN(tags);

-- Retention: drop raw metrics older than 90 days (continuous aggregates keep rollups)
SELECT add_retention_policy('metric_samples', INTERVAL '90 days');

-- Continuous aggregate: 1-hour rollups (mean, min, max, count)
CREATE MATERIALIZED VIEW metric_rollups_1h
WITH (timescaledb.continuous) AS
    SELECT
        time_bucket('1 hour', ts)   AS bucket,
        entity_id,
        metric,
        avg(value)                  AS avg_value,
        min(value)                  AS min_value,
        max(value)                  AS max_value,
        count(*)                    AS sample_count,
        (array_agg(tags))[1]        AS representative_tags
    FROM metric_samples
    GROUP BY bucket, entity_id, metric;

-- Refresh policy: refresh every 1 hour, keep 1 year of rollups
SELECT add_continuous_aggregate_policy('metric_rollups_1h',
    start_offset => INTERVAL '2 hours',
    end_offset   => INTERVAL '5 minutes',
    schedule_interval => INTERVAL '1 hour');

CREATE MATERIALIZED VIEW metric_rollups_1d
WITH (timescaledb.continuous) AS
    SELECT
        time_bucket('1 day', ts)   AS bucket,
        entity_id,
        metric,
        avg(value)                  AS avg_value,
        min(value)                  AS min_value,
        max(value)                  AS max_value,
        count(*)                    AS sample_count
    FROM metric_samples
    GROUP BY bucket, entity_id, metric;

SELECT add_continuous_aggregate_policy('metric_rollups_1d',
    start_offset => INTERVAL '2 days',
    end_offset   => INTERVAL '1 hour',
    schedule_interval => INTERVAL '1 day');

-- ─── Audit log ────────────────────────────────────────────────────────
-- Every mutating action — by the OS, by an agent, or by an operator —
-- gets an immutable audit entry. This is the "who did what when" trail
-- that the ledger doesn't fully capture (ledger records OS decisions and
-- executions; audit captures ALL API calls including reads-that-matter
-- and operator interactions).

CREATE TABLE audit_log (
    id          BIGSERIAL PRIMARY KEY,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    actor_type  TEXT NOT NULL,                 -- 'agent', 'operator', 'system', 'scheduler'
    actor_id    TEXT,                          -- entity ID of the actor (agent entity, person entity)
    action      TEXT NOT NULL,                 -- 'api.call', 'entity.create', 'policy.update',
                                               -- 'approval.decide', 'exec.request', 'deploy.trigger'
    entity_id   TEXT,                          -- entity affected (if any)
    method      TEXT,                           -- 'GET', 'POST', 'PATCH', 'DELETE', 'MCP', 'SSH'
    path        TEXT,                           -- API path or MCP tool name or SSH command
    status_code INTEGER,                        -- HTTP status or 0 for non-HTTP
    detail      JSONB DEFAULT '{}'::JSONB,     -- request body, response summary, extra context
    source_ip   TEXT,                           -- where the call came from
    correlation_id TEXT                        -- links to execution_id / signal_id for tracing
);

SELECT create_hypertable('audit_log', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_audit_actor ON audit_log(actor_type, actor_id, ts DESC);
CREATE INDEX idx_audit_entity ON audit_log(entity_id, ts DESC);
CREATE INDEX idx_audit_action ON audit_log(action, ts DESC);
CREATE INDEX idx_audit_correlation ON audit_log(correlation_id);

-- Retention: keep 1 year of audit logs
SELECT add_retention_policy('audit_log', INTERVAL '365 days');

-- ─── Event log ────────────────────────────────────────────────────────
-- Structured event stream — the "news feed" of the OS. Every significant
-- state change is an event: signal raised/resolved, execution started/completed,
-- approval requested/granted, deploy triggered, pattern validated, skill refined,
-- config changed, entity lifecycle transition. The WebSocket /api/v1/events
-- streams from this table; agents can also query it historically.

CREATE TABLE events (
    id          BIGSERIAL PRIMARY KEY,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    type        TEXT NOT NULL,                 -- 'signal.raised', 'signal.resolved',
                                               -- 'execution.started', 'execution.completed',
                                               -- 'approval.requested', 'approval.decided',
                                               -- 'deploy.triggered', 'deploy.completed',
                                               -- 'pattern.validated', 'skill.refined',
                                               -- 'entity.created', 'entity.state_changed',
                                               -- 'policy.changed', 'config.changed'
    entity_id   TEXT,                          -- primary entity involved
    severity    TEXT DEFAULT 'info',            -- info, warning, critical
    source      TEXT NOT NULL,                 -- 'scheduler', 'actuator', 'api', 'hermes',
                                               -- 'deploy', 'notifier', 'learning'
    data        JSONB DEFAULT '{}'::JSONB,     -- event-specific payload
    correlation_id TEXT                        -- links to signal/exec/approval for tracing
);

SELECT create_hypertable('events', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_events_type_ts ON events(type, ts DESC);
CREATE INDEX idx_events_entity_ts ON events(entity_id, ts DESC);
CREATE INDEX idx_events_severity_ts ON events(severity, ts DESC);
CREATE INDEX idx_events_correlation ON events(correlation_id);

-- Retention: keep 90 days of events (signals/ledger have their own tables
-- for permanent records; events are the transient feed)
SELECT add_retention_policy('events', INTERVAL '90 days');

-- ─── Agent activity log ──────────────────────────────────────────────
-- Records what the agent (Hermes) does: tool calls, reasoning, decisions,
-- token usage, latency. This is for agent behavior auditing and trend
-- analysis ("is the agent getting more efficient?").

CREATE TABLE agent_activity (
    id              BIGSERIAL PRIMARY KEY,
    ts              TIMESTAMPTZ NOT NULL DEFAULT now(),
    agent_id        TEXT NOT NULL,              -- entity ID of the agent
    session_id      TEXT,                       -- Hermes session ID
    activity_type   TEXT NOT NULL,              -- 'tool_call', 'reasoning', 'decision',
                                               -- 'mcp_query', 'ssh_command', 'escalation'
    tool_name       TEXT,                       -- MCP tool or CLI command called
    entity_id       TEXT,                       -- entity acted upon (if any)
    input_summary   TEXT,                       -- truncated input (first 500 chars)
    output_summary  TEXT,                       -- truncated output (first 500 chars)
    duration_ms     INTEGER,
    token_count     INTEGER,                    -- LLM tokens consumed (if applicable)
    success         BOOLEAN,
    correlation_id  TEXT
);

SELECT create_hypertable('agent_activity', 'ts', chunk_time_interval => INTERVAL '7 days');
CREATE INDEX idx_agent_activity_agent_ts ON agent_activity(agent_id, ts DESC);
CREATE INDEX idx_agent_activity_type_ts ON agent_activity(activity_type, ts DESC);
CREATE INDEX idx_agent_activity_entity ON agent_activity(entity_id, ts DESC);
CREATE INDEX idx_agent_activity_correlation ON agent_activity(correlation_id);

-- Retention: keep 90 days of agent activity
SELECT add_retention_policy('agent_activity', INTERVAL '90 days');

What gets captured where:

Data source Table Retention Purpose
Scheduler probes (HTTP health, disk, drift) metric_samples 90 days raw, 1 year rollups Trend analysis, anomaly detection
API response latencies metric_samples same Performance monitoring
Go runtime metrics (goroutines, mem, GC) metric_samples same OS self-monitoring
Learning model metrics (pattern confidence, skill success rate) metric_samples same Learning trend tracking
All mutating API calls audit_log 1 year Auditing, compliance, forensics
Operator actions (approval decisions, entity edits, policy changes) audit_log 1 year Operator accountability
Agent (Hermes) tool calls + reasoning agent_activity 90 days Agent behavior auditing, efficiency tracking
Signal/exec/approval/deploy state changes events 90 days Event stream (WebSocket), timeline reconstruction
Signals (permanent record) signals no expiry Signal lifecycle tracking
Ledger entries (permanent record) ledger_entries no expiry Change history
Executions (permanent record) executions no expiry Execution audit trail

Workstreams

1. Ontology definition + seed manifests (seeds/, internal/ontology/)

Before any code, finalize the ontology. The current ontology.yaml has 8 domains and 14 relationship types. The new ontology adds:

  • Layer 3 entities (cognition): signal, execution, feedback, pattern, skill, approval, classification, document, runbook
  • New relationships: triggers, produces, contributes-to, informs, guides, precedes, performs, procedure-for, learned-from
  • Lifecycle definitions for operational entities (signals, executions, patterns, skills) — not just infrastructure

Deliverables:

  • seeds/ontology.yaml — entity types, relationship types, lifecycle definitions (seeded into entity_types, relationship_types, lifecycle_defs on deploy)
  • seeds/inventory.yaml — adapted from current inventory.yaml (seeded into entities + relationships)
  • seeds/policy.yaml — adapted from current policy.yaml (seeded into risk_classes, approval_rules, autonomy_settings)
  • internal/ontology/ingest.go — idempotent seed → DB ingest

2. Database layer + migrations (migrations/, internal/db/)

  • Write the 5 migrations above
  • Set up sqlc for type-safe Go database access
  • Connection pool, migration runner
  • Graph traversal queries (blast radius, dependency chains, knowledge lookup)
  • Import script for existing signals/*.jsonl, ledger/*.jsonl → DB

3. Unified API server — Go (cmd/api/, internal/api/)

One Go binary (Gin web framework) exposing REST + MCP from the same codebase.

MCP interface (internal/api/mcp.go):

  • JSON-RPC over SSE, compatible with Hermes MCP client
  • Tools: get_host, list_services, search_knowledge, get_entity, get_relations, get_blast_radius, get_signal_history, get_ledger, get_patterns, get_skills, get_state_snapshot
  • Observability tools: query_metrics (time-series for an entity/metric), get_trend (trend analysis with slope + anomaly flags), get_audit_trail (who-did-what for an entity or actor), get_event_timeline (structured event feed for an entity or time range), get_agent_activity (what the agent did), get_health_summary (current health across the fleet with trend indicators)
  • All read from PostgreSQL

REST interface (internal/api/routes/):

  • GET /api/v1/entities — list entities (filter by type, state, domain)
  • GET /api/v1/entities/{id} — entity detail + relationships
  • POST /api/v1/entities — create entity (creates inventory entry)
  • PATCH /api/v1/entities/{id} — update entity (state transition, attributes)
  • GET /api/v1/signals — list signals (filter by state, severity, entity)
  • POST /api/v1/signals/{id}/ack — acknowledge
  • POST /api/v1/signals/{id}/resolve — resolve
  • GET /api/v1/approvals — pending approvals
  • POST /api/v1/approvals/{id}/decide — approve/deny (Authentik-gated)
  • POST /api/v1/exec — gated execution (classify → check approval → execute → feedback)
  • GET /api/v1/patterns — list patterns (filter by entity_type, action, status)
  • GET /api/v1/skills — list skills (filter by applies_to, status)
  • GET /api/v1/knowledge/{entity_id} — knowledge graph query
  • GET /api/v1/knowledge/search?q=... — search knowledge graph
  • GET /api/v1/ontology — list entity types, relationship types, lifecycles
  • POST /api/v1/ontology/entity_types — create entity type (extend the schema)
  • WS /api/v1/events — real-time stream (signals, approvals, executions, feedback)
  • Observability routes (internal/api/routes/observability.go):
    • GET /api/v1/metrics — query time-series: ?entity_id=&metric=&from=&to=&interval=
      • Returns raw samples or rollups (auto-selects 1h/1d aggregates based on range)
      • ?rollup=1h|1d|raw to force a specific resolution
      • Supports multiple metrics: ?metric=disk_usage_pct&metric=probe_latency_ms
    • GET /api/v1/metrics/{entity_id}/{metric} — single metric for one entity
      • ?from=2026-07-01T00:00:00Z&to=2026-07-06T00:00:00Z&rollup=1h
      • Returns: {entity_id, metric, samples: [{ts, avg, min, max, count}], trend: {slope, direction, anomaly}}
    • GET /api/v1/trends/{entity_id} — trend analysis for all metrics on an entity
      • Returns slope (improving/degrading/stable), recent anomalies, forecast (simple linear)
    • GET /api/v1/audit — audit log: ?actor_type=&actor_id=&entity_id=&action=&from=&to=
      • Paginated, ordered by ts DESC
      • ?correlation_id= to trace a full execution chain
    • GET /api/v1/events — historical events: ?type=&entity_id=&severity=&from=&to=
      • Same data as the WebSocket stream, but queryable historically
    • GET /api/v1/agent-activity — agent behavior log: ?agent_id=&activity_type=&entity_id=&from=&to=
      • Includes token usage, latency, success/failure per tool call
    • GET /api/v1/health — fleet health summary with trend indicators
      • Returns: {entities: [{id, type, health, trend, last_probe}], summary: {healthy, degraded, down, unknown}}
    • GET /api/v1/export — export current DB state as YAML (for DR / version control)

Policy enforcement (internal/policy/classify.go):

  • Every mutating endpoint classifies the action via the policy DB
  • Risk class → approval check → autonomy check
  • All mutations write to the ledger automatically

Auth:

  • MCP interface: no auth (internal, container-to-container)
  • REST interface: Authentik OIDC forward-auth (via Caddy) for operator endpoints
  • Internal: shared secret (Docker network)

4. Scheduler + Actuator — Go (cmd/scheduler/, internal/scheduler/, internal/actuator/)

Scheduler (Observe) — Go service with goroutines for concurrent probes:

  • HTTP health probes (concurrent, with timeouts)
  • Disk usage probes (SSH to hubris/strong)
  • Drift detection (inventory vs live state)
  • Writes signals + state snapshots to DB
  • Runs on a 10-min ticker

Actuator (Act) — Go service, the control loop:

  • Reads open signals with recommended_action
  • For each: classify via internal/policy/classify.go
    • auto-act: look up skill for (entity_type, action) → follow procedure → execute via SSH → verify → record execution → generate feedback → update patterns
    • escalate: create approval request → notify via Matrix → acknowledge signal
  • Loop-guard: check execution history per (entity, action) to cap auto-retries (SELECT ... FOR UPDATE SKIP LOCKED for concurrency safety)
  • Autonomy kill-switch: check autonomy_settings table

5. Learning engine — Go (internal/learning/)

The feedback loop that makes the agent improve over time.

Feedback recording (internal/learning/feedback.go):

  • After every execution, evaluate the outcome:
    • Did the verification command pass? → success
    • Did it fail? → failure
    • Did it partially work? → partial
    • Did something unexpected happen? → unexpected
  • Record a feedback entry with: outcome, observation (what happened vs expected), lesson (extractable insight), unexpected_side_effects

Pattern extraction (internal/learning/pattern.go):

  • Periodically scan accumulated feedback for (entity_type, action) pairs
  • When N+ executions share a similar outcome, extract a pattern:
    • "Restarting service:X typically takes 15s and succeeds"
    • "Deploying to LXC:Y via webhook has 30% failure rate, retry helps"
  • Patterns start as hypothesized, move to validated after enough evidence, then active (used by the decision classifier)
  • Confidence score = success_count / evidence_count, adjusted by recency

Skill management (internal/learning/skill.go):

  • When a pattern reaches active status with confidence > 0.7, create or refine a skill for that (entity_type, action) pair
  • Skills codify the best-known procedure (what steps to take, what to verify, expected duration, known failure modes)
  • Skills are versioned — each refinement increments the version
  • The actuator looks up skills before executing: if a skill exists, follow it; if not, use the default procedure and generate feedback for future pattern extraction

How the classifier uses learning (internal/policy/classify.go):

  • Confidence scoring now checks patterns + skills, not just raw ledger history:
    • If a pattern exists for (entity_type, action) with high confidence → boost auto-act confidence
    • If patterns show frequent failures → lower confidence, escalate
    • If a skill exists → higher confidence (proven procedure available)
  • This is the closed loop: execution → feedback → pattern → skill → classification → execution (better informed each time)

6. Knowledge graph ingestion (internal/ontology/ingest.go)

After the ontology is defined and the DB schema is in place:

  • On deploy, walk docs/ directory
  • Parse each markdown file:
    • Extract frontmatter for metadata (entity type, tags, relations)
    • Infer entity relationships from path conventions: docs/containers/105-apps.md → relationship to entity:lxc:apps
    • Extract cross-references (markdown links) → relationships
  • Create knowledge entities in the entities table (type = document, runbook, investigation, etc.) with documents / procedure-for edges to infrastructure entities
  • Idempotent — safe to re-run on every deploy

Agent access:

  • MCP tool search_knowledge(query) — full-text search on knowledge entities
  • MCP tool get_entity_knowledge(entity_id) — all docs related to an entity
  • MCP tool get_relations(entity_id) — graph traversal (blast radius, dependencies)

7. Hermes agent container (compose/hermes/)

  • Hermes Agent runtime in a Docker container, gateway mode
  • Config: hermes/config.yaml (providers, models, gateway port)
  • Persona: hermes/SOUL.md (homelab-specific)
  • Skills: hermes/skills/homelab-ops/SKILL.md — how to use the Oikos API, classify actions, request approvals, query the knowledge graph
  • Access: Oikos API via MCP (container network) + SSH keys mounted (hybrid)
  • Gateway port 8092 — workstations connect remotely
  • Hermes data volume for persistent state

8. Infisical secrets migration

Same as rev 1:

  • Stand up Infisical in the Docker stack
  • Migrate SOPS secrets (one-time decrypt + import)
  • Wire all Go services to Infisical via machine identity
  • Retire SOPS + age keys

9. Notifier — Go (cmd/notifier/, internal/notifier/)

Interface (Go):

type Notifier interface {
    SendAlert(ctx context.Context, signal Signal) error
    SendApprovalRequest(ctx context.Context, approval Approval) error
    ListenForDecisions(ctx context.Context) (<-chan ApprovalDecision, error)
}

Matrix implementation:

  • Sends alerts to @dtoro:avispero via Synapse (LXC 118)
  • Approval requests as messages with / reactions
  • Listens for reactions to record decisions
  • Pluggable — future implementations (webhook, email) register via config

10. Docker build + deploy pipeline

  • Multi-stage Dockerfiles: Go build stage → minimal runtime image (alpine or scratch)
  • docker compose build from repo root
  • Gitea webhook on push to main → deploy script on mac-mini
  • Deploy script: git pull && go build ./... && docker compose build && docker compose up -d
  • Seed ingest runs as part of the API startup (idempotent)
  • Health checks on each service

11. Ingress re-point (Caddy)

  • dtoro/caddy-conf: point mcp.hubris.network + oikos.hubris.network → mac-mini mesh IP :8090 (API)
  • Future: hermes.hubris.network → mac-mini:8092
  • DNS and public URLs unchanged

12. mac-mini host setup

  • New directory: ~/oikos-os/ — repo clone + docker compose working dir
  • Existing /opt/homelab-context/ stays untouched until cleanup phase
  • Prerequisites: Docker (or OrbStack), Go toolchain (for local dev), SSH keys
  • Cleanup (deferred): stop launchd git-sync, remove native Hermes, remove old clone

13. Decommission apps/105 (deferred)

Keep apps/105 running as fallback. Cutover checklist when ready:

  1. Verify Docker OS serves all traffic
  2. systemctl disable --now Oikos services on apps/105
  3. Remove old checkouts
  4. Update Caddy backends exclusively to mac-mini
  5. Remove/retarget Gitea webhooks

14. Observability + data capture — Go (internal/observability/)

The OS captures three classes of data, all in PostgreSQL (TimescaleDB for time-series, regular tables for audit/events):

A. Metrics (time-series)internal/observability/metrics.go:

  • Infrastructure probes: every scheduler probe writes metrics:
    • health (0=down, 1=degraded, 2=healthy) per service
    • disk_usage_pct per mount point per host
    • probe_latency_ms per probe target
    • drift_count per drift check
  • OS self-metrics: the API server exposes a /metrics endpoint and also writes to the DB:
    • api_request_count (per route, per status code)
    • api_latency_ms (p50, p99)
    • db_connections_active
    • db_query_duration_ms
    • goroutines (Go runtime)
    • gc_pause_ms
    • memory_alloc_mb
  • Learning metrics: the learning engine writes:
    • pattern_confidence per pattern ID
    • skill_success_rate per skill ID
    • auto_act_count vs escalation_count (autonomy ratio over time)
    • execution_duration_ms per (entity_type, action)
  • Agent metrics: Hermes activity is logged:
    • agent_token_count per session
    • agent_tool_call_count per session
    • agent_decision_latency_ms
  • All metrics are written to metric_samples as hypertable with 1h and 1d continuous aggregates. Raw data retained 90 days, rollups 1 year.

B. Audit log (immutable)internal/observability/audit.go:

  • Gin middleware: every mutating API call (POST/PATCH/DELETE) writes to audit_log with actor, action, entity, method, path, status, detail, source_ip
  • MCP tool calls also audited (actor_type='agent')
  • SSH commands by the actuator audited (actor_type='system', method='SSH')
  • Policy mutations get special audit entries with a hash of the before/after state
  • correlation_id propagated through the call chain (API → actuator → SSH → result) so a full execution chain can be reconstructed: "signal → classification → execution → SSH command → verification → feedback" all linked by correlation_id

C. Event streaminternal/observability/events.go:

  • Every significant state change emits an event to the events table:
    • Signal lifecycle: signal.raised, signal.acknowledged, signal.acting, signal.resolved, signal.muted
    • Execution: execution.started, execution.completed, execution.failed
    • Approval: approval.requested, approval.decided
    • Deploy: deploy.triggered, deploy.completed
    • Learning: pattern.validated, skill.refined
    • Entity: entity.created, entity.state_changed
    • Policy: policy.changed
  • The WebSocket /api/v1/events streams from this table (new events pushed to subscribers, historical events queryable via REST)
  • Events carry correlation_id for end-to-end tracing

D. Agent activityinternal/observability/agent.go:

  • Hermes MCP tool calls logged to agent_activity with tool name, entity acted upon, duration, token count, success/failure
  • SSH commands by the agent logged separately
  • Reasoning/decision audit: when the agent makes a classification decision, the reasoning is recorded (input, classification result, route, why)
  • This is the "what is the agent doing and is it getting better?" dataset

E. Structured logginginternal/observability/logging.go:

  • All services use Go's slog (structured JSON logging)
  • Every log line has: ts, level, service, msg, correlation_id (when applicable), entity_id (when applicable)
  • Logs go to stdout (Docker captures them, docker compose logs for access)
  • Debug mode: debug=true env var enables verbose probe payloads, SQL queries, classification reasoning in logs

F. Data availability for agents:

  • The agent can query its own history: "what actions have I taken on service:caddy in the last 30 days, and what were the outcomes?"
  • The agent can see trends: "is disk usage on hubris trending upward?"
  • The agent can audit: "who changed the policy for service:caddy and when?"
  • The agent can self-assess: "am I getting more efficient? (token usage trend)"
  • All through MCP tools: query_metrics, get_trend, get_audit_trail, get_event_timeline, get_agent_activity, get_health_summary

G. Future visualization plug-in points:

  • The /api/v1/metrics REST endpoint returns JSON time-series — any tool (Grafana, custom dashboard, notebook) can consume it
  • The data model is compatible with Grafana's PostgreSQL data source (time column + metric name + value + tags as labels)
  • The event stream (/api/v1/events + WebSocket) can feed a live dashboard
  • The audit log can feed a SIEM or compliance tool
  • No UI built now — the APIs are the contract; visualization plugs in later

Phasing

Phase 0 — Ontology design (no code):

  • Finalize entity types, relationship types, lifecycles
  • Write seed manifests (seeds/ontology.yaml, seeds/inventory.yaml, seeds/policy.yaml)
  • Review diagrams with operator

Phase 1 — Foundation (Go + DB):

  • Go module setup, project structure
  • PostgreSQL + TimescaleDB setup
  • Migrations 1-6 (ontology, instances, operations, learning, policy, observability)
  • Seed ingest pipeline (YAML → DB)
  • sqlc queries for core operations
  • Structured logging setup (slog)
  • Import existing signals/ledger data

Phase 2 — API (Go):

  • Gin server with REST routes
  • MCP protocol adapter (including observability tools)
  • Policy enforcement middleware
  • Audit middleware (every mutating call → audit_log)
  • Event emitter (every state change → events table)
  • Knowledge graph ingestion from docs/
  • Observability routes (metrics, trends, audit, events, health)

Phase 3 — Control loop (Go):

  • Scheduler (Observe) — probes, signals, state snapshots, metric recording
  • Actuator (Act) — classify, execute, verify, correlation ID propagation
  • Learning engine — feedback, patterns, skills, learning metrics
  • Circuit breaker per target host

Phase 4 — Agent (Hermes container):

  • Hermes Docker image, gateway config
  • Homelab skills
  • Connect from workstation, verify MCP + SSH

Phase 5 — Secrets (Infisical):

  • Stand up Infisical, migrate SOPS, wire services

Phase 6 — Deploy + cutover:

  • Docker Compose, Gitea webhook, Caddy re-point
  • End-to-end verification
  • Stop apps/105, clean up mac-mini

Reuse (logic carried over, rewritten in Go)

Existing Python What it becomes in Go
oikos/decide.py internal/policy/classify.go — same scoring logic, reads from DB + patterns
oikos/signal.py internal/scheduler/signal.go — same lifecycle, DB-backed
oikos/approve.py internal/policy/approve.go — same grant lifecycle, DB-backed
oikos/ledger.py internal/db/ — ledger_entries table + sqlc queries
oikos/policy.py + policy.yaml internal/policy/ + seeds/policy.yaml → DB tables
oikos/drift.py internal/scheduler/ — drift detection, writes signals to DB
oikos/relations.py internal/ontology/graph.go — SQL graph traversal
oikos/report.py internal/api/routes/ — report endpoints, reads from DB
mcp/server.py internal/api/mcp.go — MCP adapter on top of DB
bin/homelab logic internal/api/routes/ — same operations, REST interface
oikos/scheduler.py internal/scheduler/ — same probes, goroutines for concurrency
oikos/approve.py Matrix delivery internal/notifier/matrix.go
(new) internal/learning/ — feedback, patterns, skills (no Python equivalent)

Risks / trade-offs

  • Go rewrite — the existing Python code (~4400 lines) is replaced. The logic and design patterns carry over, but it's a full rewrite. Mitigated by the fact that the Python code is well-documented and the Go structure mirrors it.
  • Hermes in Docker — agent's world is the container. SSH access is the bridge. Hybrid approach (mounted keys now, actuator gateway later).
  • PostgreSQL as SPOF — mitigated by Docker volume persistence + automated pg_dump backups (the scheduler can do this once running).
  • Learning model cold start — no patterns/skills exist initially. The agent starts cautious (escalates everything), accumulates feedback, and gradually becomes more autonomous as patterns validate. This is by design — trust is earned.
  • Infisical bootstrapping — SOPS coexists during transition. Keep SOPS as fallback.
  • Multi-agent concurrencySELECT ... FOR UPDATE SKIP LOCKED prevents two actuator passes from acting on the same signal.
  • Ontology evolution — as the homelab changes, entity types and relationship types need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy.

Audit findings (best practices, security, performance, sanity)

A full audit was performed against the plan. Findings are organized by category and severity. High-severity items must be addressed before implementation; medium items should be addressed during the phase they belong to.

Security

ID Severity Finding Recommendation
S1 HIGH SSH private keys mounted into Hermes container — a compromised container has unrestricted SSH to all hosts Use a dedicated, restricted SSH key with command= in authorized_keys. Time-box the stopgap. Build the actuator gateway (which brokers SSH per-execution) sooner rather than later.
S2 HIGH MCP interface has no auth — any container on the Docker network has full read access to the control plane Add a shared secret or mTLS between API and Hermes. Bind MCP to a dedicated Docker network, not the default bridge. Never expose the MCP port via Caddy without auth.
S3 HIGH Policy DB is mutable via the same API it governs — a compromised API can rewrite its own approval rules (e.g. flip destructive to auto-act) Policy mutations require a meta-approval (dual-control). Add an immutable audit log of policy changes. Startup self-check: alert if policy hash differs from a known-good baseline.
S4 HIGH Learning model poisoning — flapping services or misconfigured probes can inject biased feedback to push patterns past the confidence threshold and unlock auto-act for destructive actions Require human confirmation before pattern transitions to active. Cap confidence by sample size (require N≥5 executions). Detect anomalous feedback bursts and quarantine. Never let a skill auto-promote to destructive risk class.
S5 HIGH confirmation_phrase is a weak auth primitive — if reused, stored in plaintext, or transmitted via Matrix, it's replayable Make it single-use (one phrase per approval). Store hashed. Transmit only the decision + HMAC, not the phrase. Prefer signed approval tokens.
S6 MEDIUM REST /exec and /decide endpoints rely on Authentik forward-auth only — if the API port is reachable directly on the mesh, auth disappears Enforce OIDC JWT validation in the API middleware too (defense in depth). Bind API port to localhost + Caddy only.
S7 MEDIUM No TLS between internal services — Postgres connections are plaintext on the Docker network TLS to Postgres (server cert verification). Use dedicated Docker networks per trust boundary. Document the threat model: is the Docker network trusted?
S8 MEDIUM Webhook (Gitea → mac-mini deploy) has no auth specified — anyone who can reach the endpoint can trigger arbitrary code execution via go build HMAC signature verification on the webhook. Bind listener to localhost (Gitea reaches via mesh). Run deploy script as non-root. Verify commit signatures before git pull.
S9 MEDIUM Infisical bootstrapping has a chicken-and-egg — Infisical's own master key must come from somewhere Document the bootstrap root of trust explicitly: where the master key lives (mac-mini keychain), how it's backed up, revocation path. Keep SOPS as fallback until Infisical has a tested restore-from-backup drill.
S10 MEDIUM Blast radius of a compromised container is wide — Hermes has SSH + MCP + gateway port + data volume; actuator has SSH + writes to policy/learning tables Principle of least privilege per container: only the actuator should have SSH, not Hermes. Split networks: data (PG), ops (SSH egress), front (Caddy). Use Docker user namespaces and read-only root filesystems.

Performance

ID Severity Finding Recommendation
P1 MEDIUM Recursive CTE blast_radius has no cycle guard — cycles (A→B→A) inflate work exponentially with depth Add a visited-set guard (array accumulator in the recursion). Cap max_depth at 3. Add LIMIT on the outer query. Consider a closure table for hot-path queries.
P2 MEDIUM Concurrent probes with no concurrency cap — unbounded goroutines could exhaust FDs or hammer slow targets Bounded worker pool (errgroup.SetLimit). Per-target probe timeout. Jitter to avoid thundering herd on shared backends.
P3 MEDIUM Knowledge graph re-ingestion on every deploy is wasteful if no docs changed Content-hash each doc (store hash on the entity). Skip ingestion if hash unchanged. Run in a single transaction with deferred FK checks.
P4 MEDIUM Pattern extraction frequency unspecified — either too tight (scans whole table) or too loose (patterns lag) Define cadence (hourly). Use a high-watermark on feedback.ts. Add index on feedback(ts). Process only new feedback.
P5 HIGH Table growth unaddressed — state_snapshots (every entity every 10 min) and executions/feedback grow indefinitely Partition state_snapshots and signals by month. Define TTLs (snapshots > 90 days → aggregate, raw → archive). Add a prune_* job in the scheduler.
P6 MEDIUM WebSocket /events has no backpressure — a stalled client pins a goroutine and accumulates memory Bounded channel per subscriber with drop-oldest-on-full. Max subscribers cap. Heartbeat/timeout. Document: best-effort vs. guaranteed delivery.
P7 LOW Deploy script runs go build on host AND in Docker — double build, host Go toolchain is a deploy dependency Standardize on multi-stage Docker build only. Drop host go build from deploy script. Keep host toolchain for local dev only.

Architecture

ID Severity Finding Recommendation
A1 HIGH No testing strategy — 13 workstreams, no mention of unit/integration/property tests Add a testing workstream: unit tests for classify.go, pattern.go, skill.go; integration tests with testcontainers Postgres; golden-file tests for seed ingest; property test for blast-radius (cycles, depth caps). Coverage gates per package.
A2 HIGH No observability — no structured logging, metrics, or traces. Can't inspect why the classifier escalated or how long probes took At minimum: structured JSON logs (slog) with correlation IDs per execution. A /metrics endpoint (even before full Prometheus). debug=true flag for full probe payloads. Reconsider Prometheus deferral — "who watches the watcher" requires metrics.
A3 HIGH DB backup strategy is a one-liner — the DB is the entire control-plane state (ontology, policy, inventory, learning, ledger) Define: pg_dump cadence (daily + WAL archiving for PITR), off-host storage (push to hubris or object storage, not the same mac-mini volume), encryption at rest, tested restore procedure (monthly drill), retention (30 daily + 12 monthly). This is Phase 1, not "once running."
A4 MEDIUM Seed ingest transactional safety — if API crashes mid-ingest, DB could be in partial state Run migrations + seed ingest as a distinct init container before the API starts. Wrap each seed file in a single transaction. Add a seed_version table to skip already-applied seed versions.
A5 MEDIUM Config management underspecified — unclear what's env vs. DB vs. file vs. Infisical Define a config hierarchy: defaults → file → env → Infisical (secrets only). Document each service's required config keys. Avoid putting non-secret config in Infisical.
A6 MEDIUM apps/105 fallback has no cutover safety — both stacks can't write to the same state without conflict Define cutover mode: apps/105 is read-only during coexistence. If rollback needed after new OS has been writing to Postgres, document the reconciliation plan.
A7 LOW Notifier as a separate service is over-engineered initially Either keep in-process (split later), or document the failure mode: if notifier is down, the API must retry/queue and the operator needs an alternative approval path.

Data model

ID Severity Finding Recommendation
D1 HIGH entities.id as TEXT with manual naming is fragile — renames break the ID; signal IDs require a date-string generator that's race-prone Use UUID or SERIAL for PK. Keep name + type as a unique composite for human lookup. Signal IDs: use a DB sequence.
D2 MEDIUM Ontology flexibility sacrifices type safety — attributes JSONB has no schema enforcement per entity type Add a schema JSONB column to entity_types (JSON Schema). Validate attributes against it on insert/update via a trigger or app-layer check.
D3 MEDIUM Entity type evolution has no migration story — renaming/removing an entity type cascades through many FK references Add status (active/deprecated) to entity_types. Forbid hard deletes while instances exist. Provide a merge endpoint for renames.
D4 MEDIUM Concurrent writer safety — counters on patterns (evidence_count, success_count) have read-modify-write races Use atomic UPDATE ... SET evidence_count = evidence_count + 1. Add version optimistic-lock column on skills and patterns.
D5 MEDIUM Migration rollback not addressed — only one .down.sql shown Commit to forward-only migrations (compensating migrations for rollbacks) and document it, or write and test every .down.sql.
D6 MEDIUM Data export for DR is asserted but no endpoint or format defined Add GET /api/v1/export (returns the three YAML files regenerated from DB). Test round-trip: seed → DB → export → seed → DB yields identical state.
D7 LOW relationships has no temporal data — can't answer "what depended on authentik last month?" Add valid_from/valid_to (nullable = current) for depends-on and hosts edges. Low priority but cheap to add now.

Operational

ID Severity Finding Recommendation
O1 HIGH No rollback strategy — if a deploy breaks the API (bad migration, schema bug), recovery is "revert the commit" but the migration may have already run Define: pre-deploy DB backup snapshot; migration compatibility policy (new code tolerates old schema for one deploy); tested rollback runbook per phase.
O2 HIGH "Who watches the watcher" is unresolved — if the OS is down, no alerts fire. No external/orthogonal monitor for the mac-mini or the OS containers A minimal external watchdog: a cron on apps/105 (or hubris) that curls the API /healthz every 5 min and Matrix-pings the operator directly if it fails. This must be outside the Docker stack.
O3 HIGH No backup/restore runbook — "pg_dump" is mentioned but no procedure, tested restore, or definition of what "restored" means Write docs/operations/backup-restore.md: what's backed up (DB, Infisical, Hermes volume, seed YAMLs), where, how often, how to restore each, quarterly restore drill.
O4 HIGH No disaster recovery plan — if the mac-mini dies (disk, theft, water), what's the RTO/RPO? Define RTO/RPO targets (homelab: RTO 4h, RPO 24h). Name the off-host backup target. Write the DR runbook: fresh mac-mini → install Docker → clone repo → restore Infisical → restore DB → docker compose up.
O5 MEDIUM Deploy downtime — docker compose up -d recreates containers; API has a brief gap Use docker compose up -d --no-deps for app services. Don't recreate the Postgres container on routine deploys (pin its image). Healthcheck-gated rollout.
O6 MEDIUM Health checks asserted but not specified Define per-service: API /healthz (DB ping), scheduler "last successful probe < 15min ago", notifier "last poll < 60s ago", Hermes "gateway responding". Wire into Docker healthcheck + alerting.
O7 MEDIUM mac-mini single-host failure = total outage — disk failure, macOS update reboot, Docker daemon crash Automated macOS update deferral/scheduling. Docker restart: always on all services. Monitoring heartbeat from external host. Documented cold-start runbook (what comes up first, in what order).

Missing

ID Severity Finding Recommendation
M1 HIGH No CI/CD — deploys are git-push → webhook → build. No linting, no tests before deploy, no gated merges Add CI stage (Gitea Actions): go vet, golangci-lint, go test ./..., docker build (no push). Gate the webhook on green CI. At minimum, deploy script runs go test before docker compose up.
M2 MEDIUM No rate limiting on the API — a runaway agent loop or misconfigured skill can hammer the API and DB Per-caller rate limiting (token bucket) on mutating endpoints. /exec needs a per-entity/per-action rate cap to prevent actuator loops.
M3 MEDIUM No access audit for human activity — the ledger records OS actions, but not operator API actions (entity edits, policy changes, approval decisions) Add an audit_log table for all mutating REST calls, including operator identity from OIDC.
M4 MEDIUM No circuit breaker for the actuator — if a target host is unreachable, the actuator keeps attempting SSH and generating failed executions, poisoning the learning model Circuit breaker per target: after N consecutive failures, back off (exponential) and raise a "target unreachable" signal instead of continuing to execute.
M5 MEDIUM No secret rotation story — SSH keys, Infisical machine tokens, API shared secrets need rotation policies Define rotation cadences. Document the rotation procedure for each secret class. Add a "secrets expiring" check to the scheduler.
M6 LOW No dependency/supply-chain hygiene — Go modules, Docker base images, Infisical image not pinned or scanned Pin base images by digest. Run govulncheck in CI. Periodically audit go.sum. Use distroless or scratch runtime images.
M7 LOW No documented SLOs — no quantitative success criteria (probe latency, API p99, deploy time, alert delivery) Add a small SLO table: probe interval 10min ±1min, API p99 < 200ms, deploy < 5min, alert delivery < 30s.

Top-priority items to address before implementation

  1. S1 + S10 — SSH keys in containers / wide blast radius → build the actuator gateway first, not incrementally
  2. S3 + S4 — Mutable policy DB + learning model poisoning → these compound: a compromised container can rewrite its own rules and inject feedback to unlock auto-act
  3. A1 + A2 + O2 — No tests, no observability, no external watchdog → can't safely run an autonomous agent without all three
  4. A3 + O3 + O4 — Backup/restore/DR is a one-liner for the SPOF Postgres
  5. O1 + M1 — No rollback strategy and no CI gate on deploys

Verification (end to end)

  1. Ontology: seeds/ontology.yaml ingested — SELECT * FROM entity_types shows all types across 3 layers; lifecycle definitions match the state machine diagrams.
  2. DB: docker compose up postgres — all 6 migrations applied (including TimescaleDB extension + hypertables); seed ingest populates entities + relationships from inventory.yaml.
  3. API: curl http://localhost:8090/api/v1/entities?type=service returns the fleet; MCP list_services works via the same endpoint.
  4. Scheduler: trigger a probe pass — signals in DB, state snapshots written, metrics in metric_samples (health, probe_latency_ms, disk_usage_pct).
  5. Actuator: raise a test service-down signal → actuator classifies → auto-acts (restart) or escalates (Matrix) → execution recorded → feedback generated → event emitted + audit entry written + correlation_id links the full chain.
  6. Learning: after N executions of the same (entity_type, action), a pattern appears with confidence score; after enough evidence, a skill is created. pattern_confidence metric visible in time-series.
  7. Classifier with learning: set autonomy.auto_act: reversible_low → next similar signal: classifier checks pattern confidence → auto-acts if high, escalates if low. Kill-switch (auto_act: off) → always escalates.
  8. Hermes: connect from another workstation → agent responds, queries API via MCP, can SSH to hubris. Agent activity logged to agent_activity table.
  9. Secrets: Infisical running, Go services fetch secrets, SOPS files removed.
  10. Deploy: git push → Gitea webhook → docker compose build + up -d → changes live, seed ingest syncs any YAML changes to DB. deploy.triggered + deploy.completed events in the event log.
  11. Knowledge: MCP search_knowledge("caddy") returns docs linked to entity:service:caddy.
  12. Observability: curl /api/v1/metrics?entity_id=host:hubris&metric=disk_usage_pct returns time-series with trend. curl /api/v1/audit?entity_id=service:caddy returns the full audit trail. curl /api/v1/health returns fleet summary with trend indicators. MCP get_trend("host:hubris", "disk_usage_pct") returns slope + anomaly detection.
  13. Correlation tracing: follow a correlation_id from signal → classification → execution → SSH command → verification → feedback → pattern update, all linked in the audit_log + events table.
  14. Cutover: stop apps/105, verify production traffic only from Docker OS.

Out of scope (for now)

  • Oikos Console web UI (deferred — built later on top of the API)
  • Multi-node deployment (designed for, not implemented)
  • Vector embeddings / semantic search (structured graph only for now)
  • SSH-key-signed approval requests
  • Actuator gateway pattern (Phase 2 of hybrid — start with restricted SSH key, build gateway in Phase 3 alongside the actuator)
  • Automated skill extraction via LLM (patterns extracted statistically for now; LLM-assisted skill refinement is a future enhancement)

Audit remediation — HIGH priority + architect/developer review

This section addresses all HIGH-severity audit findings and the 18 new findings from the systems architect + senior Go developer review. Each fix references the finding ID.

Schema fixes (CRITICAL/HIGH)

SA1 + SG1 — Cognition entities + broken FKs + hypertable PKs:

The core issue: the BDD models cognition objects (Signal, Execution, Feedback, Pattern, Skill, Classification) as graph entities with typed edges, but the schema only makes Signal a dual entity. executions.skill_id REFERENCES entities(id) is a broken FK (skills live in skills(id), not entities(id)). Additionally, hypertable PKs (id BIGSERIAL PRIMARY KEY) don't include the time column — TimescaleDB rejects this.

Resolution — dual entity pattern for all cognition objects:

Every cognition object gets an entities row (type = signal, execution, etc.) AND a typed table for indexed querying. The typed table's PK references entities(id). Graph relationships (triggers, produces, contributes-to) live in the relationships table, making the cognition loop traversable as a graph.

-- Fix: hypertable PKs must include the time column
CREATE TABLE audit_log (
    id          BIGSERIAL,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    -- ... columns ...
    PRIMARY KEY (id, ts)              -- was: id BIGSERIAL PRIMARY KEY
);
-- Same fix for events and agent_activity: PRIMARY KEY (id, ts)

-- Fix: executions.skill_id → references skills, not entities
-- (or if using dual-entity pattern, reference entities(id) where type='skill')
CREATE TABLE executions (
    id                SERIAL PRIMARY KEY,
    -- ... existing columns ...
    skill_id          TEXT,                           -- entity ID of skill used
    skill_version     INTEGER,                        -- SG9: snapshot version at exec time
    -- ...
);

-- Fix: skills versioning preserves history (SG9)
-- Change PK to composite so old versions remain queryable
CREATE TABLE skills (
    id                TEXT NOT NULL,                  -- 'skill-restart-service'
    version           INTEGER NOT NULL DEFAULT 1,
    name              TEXT NOT NULL,
    ts                TIMESTAMPTZ DEFAULT now(),
    procedure         TEXT NOT NULL,
    applies_to        TEXT REFERENCES entity_types(name),
    pattern_ids       TEXT[],
    status            TEXT DEFAULT 'drafted',
    success_rate      REAL,
    last_used_at      TIMESTAMPTZ,
    PRIMARY KEY (id, version)                          -- was: id TEXT PRIMARY KEY
);
CREATE TABLE skill_versions_audit (                    -- SG9: track procedure changes
    id                SERIAL PRIMARY KEY,
    skill_id          TEXT NOT NULL,
    version           INTEGER NOT NULL,
    changed_at        TIMESTAMPTZ DEFAULT now(),
    changed_by        TEXT,                            -- agent or operator entity
    diff              TEXT,                            -- diff of procedure field
    reason            TEXT
);

SA5 — Classification entity must be persisted:

Add a classifications table. Every classifier decision is recorded with the full reasoning — this is the audit trail for autonomous decisions.

CREATE TABLE classifications (
    id                SERIAL PRIMARY KEY,
    ts                TIMESTAMPTZ DEFAULT now(),
    signal_entity_id  TEXT REFERENCES entities(id),
    entity_id         TEXT REFERENCES entities(id),
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    route             TEXT NOT NULL,                  -- 'auto-act' or 'escalate'
    blast_radius      TEXT[],
    pattern_confidence REAL,
    skill_match       TEXT,                           -- skill entity ID if matched
    autonomy_check    TEXT,                           -- 'allowed' or 'blocked: <reason>'
    reasoning         JSONB NOT NULL,                 -- full decision explanation
    correlation_id    TEXT
);
CREATE INDEX idx_class_signal ON classifications(signal_entity_id);
CREATE INDEX idx_class_entity ON classifications(entity_id);

SG2 — CAGG array_agg unsupported:

Drop representative_tags from the 1h continuous aggregate. Query tags from raw data when needed.

CREATE MATERIALIZED VIEW metric_rollups_1h
WITH (timescaledb.continuous) AS
    SELECT
        time_bucket('1 hour', ts)   AS bucket,
        entity_id,
        metric,
        avg(value)                  AS avg_value,
        min(value)                  AS min_value,
        max(value)                  AS max_value,
        count(*)                    AS sample_count
        -- removed: (array_agg(tags))[1] AS representative_tags
    FROM metric_samples
    GROUP BY bucket, entity_id, metric;

SG3 — TimescaleDB functions not idempotent:

Use if_not_exists => TRUE and exception guards:

SELECT create_hypertable('metric_samples', 'ts',
    chunk_time_interval => INTERVAL '7 days',
    if_not_exists => TRUE);

DO $$ BEGIN
    PERFORM add_retention_policy('metric_samples', INTERVAL '90 days');
EXCEPTION WHEN OTHERS THEN NULL;
END $$;

Ontology fixes (HIGH)

SA2 — Layer boundary "cognition creates governance" arrow:

Remove the unsupported arrow. The learning engine does NOT write to governance (policy/autonomy) tables. Instead, validated patterns propose autonomy changes as approval-request entities — the operator must accept. This keeps governance authoritative and prevents the learning-poisoning vector (S4).

Also rename the cognition subgraph node from "Governance" to "Approvals" to avoid the layer-name collision.

SA3 — Missing entity types (Person, Agent, IdentityProvider):

Add BDD definitions and entity_type seeds:

Person: matrix_id, oidc_sub
Agent: provider, model, gateway_port
IdentityProvider: issuer, client_id, auth_mode (oidc/forward-auth)

These are first-class governance-layer entities. audit_log.actor_id now resolves to a real entity.

SA4 — Lifecycle dead-ends and missing transitions:

Updated infrastructure lifecycle (add failed state):

planned → provisioning → active → migrating → deprecated → destroyed
                    ↘ failed ↗          ↘ failed ↗
failed → active (recovery) | failed → deprecated (write-off)
planned → destroyed (cancel)
deprecated → active (un-deprecate if replacement fails)

Updated signal lifecycle (add terminal states):

raised → acknowledged → acting → resolved | failed
acknowledged → resolved (manual resolve without acting)
acknowledged → muted
acting → failed (permanent failure, terminal — needs operator)
failed → acknowledged (operator retries)

Updated execution lifecycle (add cancellation + rollback failure):

approved → expired (TTL ran out)
executing → cancelled (operator abort)
failed → rolled_back | rollback_failed
timed_out → verifying (check if the command actually completed despite timeout)

Updated pattern lifecycle:

hypothesized → validated → active → deprecated
hypothesized → invalidated (disproven, terminal)
active → invalidated (new evidence contradicts)

Updated skill lifecycle:

drafted → tested → active → refined → active (new version)
drafted → deprecated (abandoned)
tested → failed → drafted (back to drawing board)
active → deprecated (superseded or unsafe)

Approval lifecycle (new diagram — was missing):

pending → approved | denied | expired
approved → revoked (operator changes mind before execution)

SA6 — recommended_action data source:

Move recommended_action from signals to classifications. The probe raises a signal (kind, severity, evidence). The classifier populates the classification with the recommended action based on signal kind + skill lookup. The actuator reads the classification, not the signal, for the action to take.

SA8 — Missing entities (Cluster, ComposeStack, ManagedHost):

Add to ontology:

  • Cluster entity (software domain) — ProxmoxHost member-of Cluster
  • ComposeStack entity (software domain) — DockerContainer part-of ComposeStack
  • Add provider, control_level (full/partial/none) to StandaloneServer

Security remediation (HIGH)

S1 — SSH keys in containers:

Dual approach (operator decision):

  1. Phase 1-2 (immediate): Restricted SSH key — dedicated key pair with command="..." and from="..." restrictions in authorized_keys on hubris/ strong. The key can only run specific commands (pct, qm, df, systemctl status), not arbitrary shells. Mounted read-only into the actuator container only (not Hermes).
  2. Phase 3 (actuator build): Full actuator gateway — the API's /exec endpoint brokers all SSH. The actuator holds the keys, executes per-action, logs every command. Hermes never touches SSH.

S2 — MCP auth:

  • Shared secret between API and Hermes (HMAC-signed requests)
  • MCP bound to a dedicated Docker network (not the default bridge)
  • Never exposed via Caddy without auth

S3 — Policy DB mutability:

  • Policy mutations (risk_classes, approval_rules, autonomy_settings) require a meta-approval: the operator must approve the policy change itself (dual-control)
  • Immutable audit log of all policy changes with before/after hash
  • Startup self-check: compute policy hash, alert if differs from last-known-good

S4 — Learning model poisoning:

  • Pattern transitions to active require operator confirmation (PATCH /api/v1/ patterns/{id} with status=active — policy-gated as config_mutation)
  • Confidence capped by sample size: confidence = min(raw_confidence, N/5) where N = evidence_count (requires N≥5 for confidence > 0.2)
  • Anomaly detection: if >10 identical-outcome feedback entries arrive within 1 hour for the same (entity_type, action), quarantine the pattern for review
  • Skills can never auto-promote to destructive risk class — always escalate

S5 — confirmation_phrase replacement:

  • Replace with single-use signed approval tokens
  • Token = HMAC(approval_id + entity_id + action + risk_class + nonce, shared_secret)
  • Stored hashed in approvals table
  • Transmitted via Matrix as the approval ID + decision; token verified server-side

SA10 — Gateway + Caddy trust boundary:

  • Port 8092 (Hermes gateway): mTLS or token auth. Mesh membership is the network boundary; gateway auth is the application boundary.
  • Caddy is an explicit trust root. API validates OIDC JWTs in middleware (not just trusting Caddy headers). Documented: compromising Caddy ≠ compromising the API.

Operational remediation (HIGH)

A3 + O3 + O4 — Backup, restore, DR:

Backup strategy:

  • Daily pg_dump (compressed, custom format) + WAL archiving for PITR
  • Off-host storage: Proton Drive (cloud object storage, encrypted at rest)
  • Push via rclone from the scheduler container (already have rclone LXC in the fleet — reuse credentials)
  • Retention: 30 daily + 12 monthly snapshots
  • Infisical backup: Infisical has its own backup mechanism; also export secrets to an encrypted SOPS file as a fallback (chicken-and-egg: keep one age key for this purpose)
  • Hermes volume: backed up with pg_dump of agent_activity + session data

Restore procedure (docs/operations/backup-restore.md):

  1. Restore Postgres: pg_restore -d oikos < dump.psql
  2. Verify seed ingest matches (run GET /api/v1/export and diff against seed YAML)
  3. Restore Infisical from its backup
  4. docker compose up -d
  5. Monthly restore drill (scheduled, automated, alert if restore fails)

DR plan:

  • RTO: 4 hours (fresh machine → Docker → restore → running)
  • RPO: 24 hours (last daily backup)
  • Cold-start runbook: install Docker → clone repo → restore Infisical → restore DB → docker compose up -d → verify health
  • Off-host backup target: Proton Drive (encrypted, offsite)

O1 — Rollback strategy:

  • Forward-only migrations (no down.sql beyond development). Compensating migrations for production rollbacks.
  • Pre-deploy DB backup: the deploy script runs pg_dump before docker compose up -d
  • Migration compatibility: new code must tolerate old schema for one deploy window (additive migrations only — new columns nullable, new tables optional)
  • Rollback runbook: revert git commit → pg_restore from pre-deploy backup → docker compose up -d with old image

O2 — External watchdog:

  • Cron job on apps/105 (outside the Docker stack): curl -sf http://mac-mini:8090/healthz || curl -X POST matrix-webhook ...
  • Runs every 5 minutes
  • Alerts operator directly via Matrix if the API is unreachable
  • Also checks: Docker daemon running (docker info), Postgres accepting connections (pg_isready)

M1 — CI/CD:

  • Gitea Actions (or simple webhook + script):
    • go vet ./...
    • golangci-lint run
    • go test ./... -race -cover
    • docker build (no push — just verify it builds)
  • Webhook deploy gated on green CI
  • Deploy script runs go test as a final safety check before docker compose up

Architecture remediation (HIGH)

A1 — Testing strategy:

Add testing workstream with specific tests:

Package Test type What to test
internal/policy/ Unit Classifier scoring: risk × blast × confidence. Table-driven: every (risk_class, blast_radius, confidence) combination. Edge: unknown action, ambiguous entity.
internal/policy/ Unit Approval lifecycle: token issue, verify, single-use enforcement, TTL expiry.
internal/learning/ Unit Pattern confidence calculation. Skill versioning. Feedback → pattern extraction.
internal/ontology/ Unit Lifecycle transition validation: every legal transition succeeds, every illegal one fails. Graph traversal (mock relationships).
internal/db/ Integration testcontainers Postgres: seed ingest idempotency, blast_radius CTE (with cycles), hypertable insert + query, continuous aggregate refresh.
internal/api/ Integration testcontainers: REST routes return correct status codes, MCP tools return expected shapes, audit middleware records entries, pagination works.
internal/actuator/ Integration Mock SSH: execution → verification → feedback → pattern update. Loop-guard prevents retry storms. Circuit breaker trips after N failures.
migrations/ Property Every migration is forward-only. blast_radius returns correct results on cyclic graphs. Hypertable retention doesn't drop data younger than threshold.
internal/observability/ Unit Correlation ID propagation through context. Event emitter transactional with state change. Metric recording.

Coverage gate: ≥80% on internal/policy/ and internal/learning/ (the autonomy- granting code). ≥60% on everything else.

A2 — Observability: Already addressed (Migration 6, Workstream 14). Update audit status to resolved.

SG4 — Graceful shutdown:

Every cmd/*/main.go implements:

  1. signal.NotifyContext(ctx, syscall.SIGTERM, syscall.SIGINT)
  2. Context propagated to all long-running loops and HTTP server
  3. Shutdown sequence: stop accepting new work → wait for in-flight (30s deadline) → for actuator: if execution in-flight, mark failed with "shutdown interrupted" + emit feedback → close DB pool
  4. docker-compose.yml: stop_grace_period: 30s on actuator, stop_signal: SIGTERM on all services

SG5 — Entity-level concurrency:

-- Per-entity advisory lock during execution (prevents concurrent actions
-- on the same target entity, e.g., restart + deploy on the same service)
SELECT pg_advisory_xact_lock(hashtext($1));  -- $1 = target_entity_id
-- ... execute, verify, feedback ...
-- lock released on transaction commit/rollback

SG6 — Domain layer:

Add internal/domain/ package:

internal/domain/
├── entity.go        # Entity, EntityType, Relationship domain types
├── signal.go        # Signal domain type + lifecycle transition logic
├── execution.go     # Execution domain type + state machine
├── classification.go # Classification domain type
├── pattern.go       # Pattern domain type + confidence calculation
├── skill.go         # Skill domain type + versioning
├── approval.go      # Approval domain type + token verification
└── errors.go        # Sentinel errors: ErrNotFound, ErrInvalidTransition,
                     # ErrApprovalRequired, ErrAutonomyBlocked, ErrConflict

DB ↔ domain mapping in internal/db/ (repository pattern). API handlers accept/ return domain types. sqlc models never escape internal/db/.

SG11 — Error handling:

// internal/domain/errors.go
var (
    ErrNotFound          = errors.New("entity not found")
    ErrInvalidTransition = errors.New("invalid lifecycle transition")
    ErrApprovalRequired  = errors.New("operator approval required")
    ErrAutonomyBlocked   = errors.New("autonomy policy blocks this action")
    ErrConflict          = errors.New("concurrent modification conflict")
    ErrCircuitOpen       = errors.New("circuit breaker open for target")
)

HTTP mapping middleware: ErrNotFound → 404, ErrInvalidTransition → 409, ErrApprovalRequired → 403, ErrAutonomyBlocked → 403, ErrConflict → 409, ErrCircuitOpen → 503.

SSH error classification in internal/actuator/execute.go:

  • Network unreachable → retryable, circuit breaker
  • Auth failure → fatal, alert operator
  • Command exit non-zero → execution failed, feedback
  • Command timeout → timed_out, feedback

DB retry: serialization failures (SQLSTATE 40001, 40P01) → retry with exponential backoff (max 3 retries).

Go implementation fixes (MEDIUM)

SA9 — TimescaleDB Docker image + migration runner:

  • Image: timescale/timescaledb:2.x-pg16 (not postgres:16)
  • Migrations run in a one-shot init container (compose/migrate/Dockerfile) with a dedicated DB user that has DDL but no runtime data privileges
  • API's DB user gets DML only (least privilege)
  • golang-migrate Go API with //go:embed migrations/*.up.sql

SG7 — Pattern/skill management endpoints:

  • PATCH /api/v1/patterns/{id} — state transition (validate/invalidate/deprecate), policy-gated as config_mutation, audit-logged
  • PATCH /api/v1/skills/{id} — state transition + version pin, same gating
  • This is the operator's manual safety valve for learning-model issues (S4)

SG8 — WebSocket push mechanism:

  • In-process event bus (Go channel pub/sub) for events written by the API itself (zero-latency push to WebSocket subscribers)
  • Postgres LISTEN/NOTIFY for events written by other services (scheduler, actuator) — trigger on events table fires NOTIFY after commit
  • Both feed the WebSocket handler

SG10 — Transactional event emission:

  • Event + audit entries written in the same DB transaction as the state change
  • If transaction rolls back, events are discarded (never emitted)
  • LISTEN/NOTIFY fires after commit — subscribers only see committed events

SG13 — Context-aware SSH:

func runSSH(ctx context.Context, client *ssh.Client, cmd string) ([]byte, error) {
    session, err := client.NewSession()
    if err != nil { return nil, err }
    defer session.Close()
    type result struct { out []byte; err error }
    ch := make(chan result, 1)
    go func() { out, err := session.CombinedOutput(cmd); ch <- result{out, err} }()
    select {
    case r := <-ch: return r.out, r.err
    case <-ctx.Done():
        session.Close()  // unblocks CombinedOutput
        client.Close()
        return nil, ctx.Err()
    }
}

SG14 — Connection pool sizing:

  • API: 15 connections, scheduler: 5, actuator: 5, learning: 3 = 28 total
  • Postgres max_connections set to 80
  • Monitor db_connections_active, alert if >80% of pool

SG15 — RESTful exec endpoint:

  • POST /api/v1/executions (was POST /api/v1/exec) — creates an execution resource. Handler classifies, checks approval, creates execution in proposed state.
  • GET /api/v1/executions/{id} — status
  • POST /api/v1/executions/{id}/cancel — cancellation

SG16 — Pagination:

  • All hypertable-backed endpoints: cursor-based (?cursor=<ts>&limit=50)
  • entities and other small tables: keyset pagination (?after=<id>&limit=50)
  • Default limit: 50, max: 200
  • MCP tools support limit parameter

SG17 — Go tooling:

  • sqlc.yaml added to repo layout
  • Module path: github.com/dtoro/oikos
  • CGO_ENABLED=0 in Dockerfiles, pgx (pure Go), gcr.io/distroless/static runtime image
  • Migrations embedded with //go:embed

SG18 — Health/metrics bypass auth:

  • /healthz and /metrics on a separate Gin router group, no auth, no audit
  • /healthz: SELECT 1 against DB
  • /metrics: internal-only (not exposed via Caddy), or token-protected

SA7 — Notifier decoupling:

  • API writes approvals row (status=pending) + emits approval.requested event
  • Notifier polls pending approvals, sends to Matrix, writes decision directly to approvals table (has DB access, not API access)
  • API polls approvals.status
  • No service-to-service calls in either direction — DB is the rendezvous point
  • Pending approvals survive Notifier restart

SA9 + SA10 — Threat model documentation:

  • Caddy is an explicit trust root (compromise = API compromise, mitigated by JWT validation in API middleware)
  • Mesh membership is the network boundary for Hermes gateway
  • Docker network is the trust boundary for internal services (mTLS between API and Hermes, TLS to Postgres)
  • The actuator is the only container with SSH egress (not Hermes, not the API)

Updated phasing (incorporating remediation)

Phase 0 — Ontology design (no code):

  • Finalize entity types (including Person, Agent, IdentityProvider, Cluster, ComposeStack) + relationship types + lifecycles (all with terminal states)
  • Write seed manifests
  • Review diagrams + lifecycle completeness with operator

Phase 1 — Foundation (Go + DB):

  • Go module setup (github.com/dtoro/oikos), project structure with domain layer
  • PostgreSQL + TimescaleDB (timescale/timescaledb:2.x-pg16)
  • Migrations 1-6 (with fixed hypertable PKs, idempotent TimescaleDB calls, dual entity pattern for cognition objects, classifications table)
  • Migration runner as init container (DDL-only DB user)
  • Seed ingest pipeline (transactional, init container)
  • sqlc + domain layer (repository pattern)
  • Structured logging (slog) + error sentinels + HTTP error mapping
  • Testing foundation: testcontainers setup, unit test framework, coverage gates
  • Backup setup: daily pg_dump + rclone push to Proton Drive, WAL archiving
  • External watchdog: cron on apps/105

Phase 2 — API (Go):

  • Gin server with REST routes (resource-oriented, paginated, /executions not /exec)
  • MCP protocol adapter (shared-secret auth, dedicated network)
  • Policy enforcement middleware (OIDC JWT validation, not just Caddy headers)
  • Audit middleware (transactional with state changes)
  • Event emitter (transactional, in-process bus + LISTEN/NOTIFY)
  • Observability routes (metrics, trends, audit, events, health, agent-activity)
  • WebSocket (in-process bus + LISTEN/NOTIFY, bounded channels, backpressure)
  • Pattern/skill management endpoints (operator override for learning model)
  • Domain layer fully fleshed out
  • CI: Gitea Actions (go vet, golangci-lint, go test -race -cover, docker build)

Phase 3 — Control loop (Go):

  • Scheduler (Observe) — probes, signals, state snapshots, metric recording
  • Actuator (Act) — classify, execute, verify, correlation ID propagation
    • Restricted SSH key (command= in authorized_keys, actuator-only)
    • Entity-level advisory locks (pg_advisory_xact_lock)
    • Context-aware SSH (context cancellation, hard timeout)
    • Circuit breaker per target host
    • Graceful shutdown (SIGTERM, in-flight protection, stop_grace_period: 30s)
  • Learning engine — feedback, patterns, skills, learning metrics
    • Pattern activation requires operator confirmation (PATCH endpoint)
    • Confidence capped by sample size (N≥5)
    • Anomaly detection for feedback bursts
  • Approval tokens (single-use HMAC, not confirmation phrases)
  • Policy meta-approval (dual-control for policy mutations)
  • Notifier (DB rendezvous, no service-to-service calls, Matrix impl)

Phase 4 — Agent (Hermes container):

  • Hermes Docker image, gateway config (mTLS on port 8092)
  • Homelab skills
  • Agent activity logging
  • Connect from workstation, verify MCP (shared-secret auth) + SSH via actuator

Phase 5 — Secrets (Infisical):

  • Stand up Infisical, migrate SOPS, wire services
  • Infisical backup + SOPS fallback (one age key kept for DR)

Phase 6 — Deploy + cutover:

  • Docker Compose (timescale image, init containers, pool sizing, stop_grace_period)
  • Gitea webhook (HMAC auth, non-root deploy user, CI-gated)
  • Caddy re-point + JWT validation in API
  • End-to-end verification (14 checks)
  • Stop apps/105, clean up mac-mini
  • Restore drill (monthly, automated)