Files
oikos/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md
dtoro 18cb79caf9 oikos phase 0: ontology + inventory + policy seeds, OpenAPI contract, ADRs
- seeds/ontology.yaml: 59 entity types (5 abstract, is-a hierarchy), 46
  relationship types with cardinality, 6 lifecycles with terminal states
  and named precondition checks
- seeds/inventory.yaml: 110 entities / 142 relationships translated from
  legacy inventory.yaml (fleet, services, ingress, storage, governance,
  archaeology); thin spots marked for backfill
- seeds/policy.yaml: 4 risk classes, 27 approval rules (hierarchy-aware,
  per-entity overrides), autonomy kill-switch off (cold start)
- api/openapi.yaml: full v1 REST contract (40 paths), RFC 9457 errors,
  cursor pagination, idempotency, ETag/If-Match, scopes; redocly-clean
- docs/adr/0001-0010: initial architecture decision records
- scripts/validate-seeds.py: Phase 0 gate — hierarchy, lifecycles,
  endpoints, cardinality, policy cross-refs (0 errors)
- plan: layer CHECK gains 'meta' (root type), cardinality gains
  'many-to-one'

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 00:17:15 +02:00

77 KiB
Raw Permalink Blame History

Plan: Oikos — Docker-based agentic homelab OS on mac-mini

Status: Planned (2026-07-06, rev 3) — supersedes rev 2. Rev 3 consolidates the rev-2 audit + remediation layers into one self-consistent spec (no more "read the migration, then read the fix section") and closes newly found gaps. This document is the single source of truth for implementation; an agent should be able to implement phase by phase from this file alone.

Rev 3 changelog

Consolidation:

  • All rev-2 HIGH/MEDIUM remediations (S1S10, SA1SA10, SG1SG18, A1A7, O1O7, D1D7, P1P7, M1M7) are now merged inline. Appendix A maps every finding ID to where it is resolved in this document.

New in rev 3 (gaps found in external review):

  • R3-1 Ontology inheritance in the meta-schema. The BDD uses generalization heavily (ComputeEntity <|-- Machine <|-- ProxmoxHost), and relationships hang off abstract types (ComputeEntity provides Service), but the rev-2 entity_types table was flat. Added parent_type + is_abstract; relationship endpoint validation walks the type hierarchy.
  • R3-2 Contract-first API. api/openapi.yaml (OpenAPI 3.1) is the source of truth; Go server stubs generated with oapi-codegen (chi router, strict server). Future UIs get a generated TypeScript client. Gin dropped.
  • R3-3 API semantics for machines and UIs. RFC 9457 problem+json errors, uniform list envelope, cursor pagination, Idempotency-Key on unsafe POSTs, ETag/If-Match optimistic concurrency, role scopes (operator/viewer/agent), CORS policy, and a /graph endpoint for UI visualization.
  • R3-4 Single binary, role subcommands. One cmd/oikos binary (oikos api| scheduler|notifier|all), one Docker image, roles selected by compose command (Loki/Temporal pattern). Guarantees version consistency; trivial local dev.
  • R3-5 IDs. UUIDv7 primary keys (app-generated, time-ordered) + unique human slug (host:hubris). API accepts either. Resolves D1 properly.
  • R3-6 P5 actually fixed. state_snapshots (unbounded growth) replaced by a single-row-per-entity entity_status table; health history lives in metric_samples (already retained/rolled up).
  • R3-7 Checks as data. Probe definitions (check_defs) live in the DB and are ontology-attached — adding a probe is an API call, not a code deploy.
  • R3-8 Signal dedup + flap suppression + maintenance mode. Partial unique index guarantees one open signal per (entity, kind); occurrence counting; hold-down on flapping; maintenance_until on entities suppresses signals/auto-act.
  • R3-9 Executable skill format. skills.procedure is a JSON-schema-validated structure (steps/verify/rollback/params) the actuator can run deterministically; markdown is rendered from it for humans.
  • R3-10 MCP modernized. Official github.com/modelcontextprotocol/go-sdk, Streamable HTTP transport (SSE-only transport is deprecated in the MCP spec), static bearer token auth. Tools are thin wrappers over the same service layer as REST — one behavior, two protocols.
  • R3-11 Ledger simplified. ledger_entries table dropped; ledger is a SQL view over executions ⋈ classifications ⋈ approvals. One less write path to keep consistent; audit_log already covers operator mutations.
  • R3-12 Release + rollback made concrete. Images tagged with git SHA, last 5 kept; rollback = redeploy previous tag + pg_restore of pre-deploy dump.
  • R3-13 macOS deployment realities + dual-path networking. Docker-in-VM (OrbStack), no host networking, sleep/auto-restart settings, launch-at-login. Outbound SSH goes direct over the LAN; inbound is mesh-primary with a LAN break-glass binding for the API (rev 3.1, operator decision). Previously silent.
  • R3-14 SSE event stream. Server→client push is all we need; SSE is simpler than WebSocket through Caddy and for browser UIs. WebSocket deferred.
  • R3-15 Phases now carry acceptance criteria ("done when" + verification commands) so an implementing agent knows when to stop.
  • R3-16 ADRs. docs/adr/ with MADR template; the Decisions table below seeds the initial ADRs. Future architecture changes are recorded, not re-litigated.

Vision

Convert this repo into a Docker-based agentic homelab OS written in Go. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors:

  • Operator (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix.
  • Agent — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and learns from every action.

All OS services run in Docker on mac-mini, deployed by git push (Gitea webhook → image build → restart). Designed for mac-mini now with a path to multi-node later (see "Multi-node path").

Decisions

Question Decision
Repo structure One repo, reorganized (see Repo layout).
Language Go — compiled, type-safe, small containers, goroutines for concurrent probes.
Packaging Single binary oikos with role subcommands; one Docker image (R3-4).
API style OpenAPI-firstapi/openapi.yaml is the contract; oapi-codegen + chi; RFC 9457 errors (R3-2/3).
Hermes runtime Docker container, gateway mode.
Agent → homelab access Hybrid — restricted SSH key in the actuator only now; full actuator gateway in Phase 3. Hermes never holds SSH keys.
Data storage PostgreSQL 16 + TimescaleDB. DB is the runtime source of truth.
Config (inventory/ontology/policy) DB-native; YAML files are seed manifests (bootstrap + DR) with round-trip export.
IDs UUIDv7 PK + unique slug for humans/API (R3-5).
Operator interface Remote Hermes + Matrix now; UIs later on top of the API.
Deploy Git push → CI green → Gitea webhook → SHA-tagged image build → compose up.
bin/homelab CLI Thin Go client generated from the OpenAPI spec.
Secrets Migrate SOPS+age → Infisical (one age key kept for DR fallback).
Notifications Matrix now, behind a Notifier interface; DB is the rendezvous (no service-to-service calls).
MCP Same binary/service layer as REST; official Go SDK, Streamable HTTP (R3-10).
Event stream SSE (/api/v1/events/stream); WebSocket deferred (R3-14).
Feedback loop Agent learns from execution; pattern activation and any autonomy expansion require operator approval (anti-poisoning).
Ontology Developed first; meta-schema supports inheritance + abstract types (R3-1).
apps/105 Keeps running (read-only toward shared state) as fallback until cutover.
ADRs docs/adr/ (MADR format); this table seeds ADR-0001…0010 (R3-16).

Ontology — the systems model

The ontology defines what exists, how things connect, how they change over time, and how the OS learns. It is stored in the database (entity types, relationship types, lifecycle definitions). YAML seeds bootstrap it; after that the DB is authoritative and editable via API.

Design principles

  1. Three layers — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's behavior + learning). Dependencies flow downward: Cognition depends on Governance depends on Infrastructure.
  2. Everything is an entity — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, executions, skills) are first-class entities.
  3. Typed relationships with cardinality — edges carry semantics and are queryable (blast radius, dependency chains, knowledge lookup).
  4. Inheritance is part of the model — entity types form an is-a hierarchy with abstract types (ComputeEntity); relationship endpoint constraints may name abstract types and validation walks the hierarchy (R3-1).
  5. Lifecycles are state machines — every entity type has a lifecycle with explicit terminal states and (named, code-implemented) transition preconditions.
  6. Policy attaches to the ontology — risk classes and approval rules link to entity types and actions.
  7. Learning is modeled but never self-authorizing — executions → feedback → patterns → skills is an explicit, queryable graph, but the learning engine only proposes governance changes; the operator approves them (S4/SA2).

Layer map

graph TB
    cognition -- "observes, acts on, learns about" --> infra
    governance -- "governs access to" --> infra
    governance -- "constrains" --> cognition
    cognition -. "proposes changes (operator approves)" .-> governance

    subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"]
        direction LR
        OBS["Observation\nsignal, check, entity-status"]
        DEC["Decision\nclassification"]
        ACT["Action\nexecution, verification"]
        APPR["Approvals\napproval-request, approval-decision"]
        KNOW["Knowledge\ndocument, runbook"]
        LEARN["Learning\nfeedback, pattern, skill"]
    end

    subgraph governance["Layer 2 — Governance (who controls what)"]
        direction LR
        IDENT["Identity\nperson, agent, identity-provider"]
        SEC["Secrets\nsecret, key, access-grant"]
        POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
    end

    subgraph infra["Layer 1 — Infrastructure (the managed world)"]
        direction LR
        PHYS["Physical\nsite, machine, ups, sensor"]
        COMP["Compute\nmachine, vm, container\n(lxc, docker)"]
        NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
        STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
        SOFT["Software\nservice, application, cluster,\ncompose-stack, config-repo,\ndeploy-pipeline"]
    end

Note the dashed arrow: the learning engine cannot write to governance tables. A validated pattern that would expand autonomy becomes an approval request; only an operator decision changes policy (SA2, S4).

Block definition diagrams (SysML BDD)

Conventions: «abstract» = cannot be instantiated; <|-- generalization; *-- composition; o-- aggregation; --> association; multiplicities as labeled.

Infrastructure — compute, storage, network:

classDiagram
    class ComputeEntity {
        <<abstract>>
        +state lifecycle
        +attributes jsonb
    }
    class Machine { +cpu_arch +ram_gb }
    class VirtualMachine { +vcpus +memory_mb +disk_gb }
    class Container { <<abstract>> +runtime }
    class LXC { +pve_id +rootfs }
    class DockerContainer { +image +compose_stack }
    class ProxmoxHost { +pve_version +cluster_member }
    class StandaloneServer { +hypervisor +provider +control_level }
    class Workstation { +os +user }
    class Appliance { +vendor +model }
    class Hypervisor { +type +version }
    class Cluster { +quorum +members }
    class ComposeStack { +path +services }

    ComputeEntity <|-- Machine
    ComputeEntity <|-- VirtualMachine
    ComputeEntity <|-- Container
    Machine <|-- ProxmoxHost
    Machine <|-- StandaloneServer
    Machine <|-- Workstation
    Machine <|-- Appliance
    Container <|-- LXC
    Container <|-- DockerContainer

    Machine "1" *-- "0..1" Hypervisor : runs
    Hypervisor "1" o-- "0..*" VirtualMachine : hosts
    Hypervisor "1" o-- "0..*" Container : hosts
    ProxmoxHost "0..*" --> "0..1" Cluster : member-of
    DockerContainer "0..*" --> "0..1" ComposeStack : part-of

    class StoragePool { +type lvm, zfs, nfs +capacity_gb }
    class Volume { +name +size_gb }
    class Mount { +mount_point +options }

    StoragePool "1" *-- "0..*" Volume : contains
    ComputeEntity "1" o-- "0..*" Mount : has
    Mount "0..*" --> "1" Volume : mounts

    class NetworkInterface { +mac +ip }
    class Network { <<abstract>> }
    class LAN { +subnet }
    class Mesh { +provider }
    class VLAN { +tag }

    ComputeEntity "1" *-- "0..*" NetworkInterface : has
    NetworkInterface "0..*" --> "1" Network : connects-to
    Network <|-- LAN
    Network <|-- Mesh
    Network <|-- VLAN

Software + services:

classDiagram
    class Service { +port +health_url +risk_notes }
    class Application { +version +config }
    class ConfigRepo { +url +branch }
    class DeployPipeline { +trigger +target_path }
    class IngressRoute { +pattern +upstream }
    class Certificate { +issuer +expires }
    class DNSZone { +zone }
    class DNSRecord { +name +record_type +value }

    ComputeEntity "1" o-- "0..*" Service : provides
    Service "1" *-- "0..*" Application : runs
    Service "0..1" --> "0..1" ConfigRepo : configured-by
    DeployPipeline "0..*" --> "1" Service : deploys-to
    IngressRoute "0..*" --> "1" Service : routes-to
    IngressRoute "0..*" --> "0..1" Certificate : secured-by
    IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by
    Service "0..*" --> "0..*" Service : depends-on
    DNSZone "1" *-- "0..*" DNSRecord : contains
    DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to

Governance — identity (new in rev 2 remediation, kept):

classDiagram
    class Person { +matrix_id +oidc_sub }
    class Agent { +provider +model +gateway_port }
    class IdentityProvider { +issuer +client_id +auth_mode }
    class Secret { +path +rotation_days }
    class AccessGrant { +scope +expires }

    Person "0..1" --> "0..*" Agent : owns
    IdentityProvider "1" o-- "0..*" Person : authenticates
    AccessGrant "0..*" --> "1" Secret : grants
    Agent "0..*" --> "0..*" AccessGrant : holds

Cognition — operations + learning:

classDiagram
    class CheckDef { +kind +config +interval }
    class Signal { +kind +severity +state +evidence +occurrences }
    class Classification { +risk +route +reasoning }
    class Execution { +status +result +duration_ms +verified }
    class Feedback { +outcome +observation +lesson }
    class Pattern { +confidence +evidence_count +status }
    class Skill { +procedure +version +success_rate }
    class Approval { +status +ttl }
    class Document { +title +content +source_path }
    class Runbook { +steps +risk_class +verification }

    CheckDef "0..*" --> "1" Entity : checks
    CheckDef "1" o-- "0..*" Signal : raises
    Signal "1" --> "0..*" Classification : classified-by
    Classification "1" --> "0..1" Execution : precedes
    Execution "1" *-- "0..1" Feedback : produces
    Feedback "0..*" --> "0..*" Pattern : contributes-to
    Pattern "0..*" --> "0..1" Skill : informs
    Skill "0..1" --> "0..*" Classification : guides
    Execution "0..1" --> "0..1" Approval : requires
    Agent "0..*" --> "0..*" Execution : performs
    Person "0..1" --> "0..*" Approval : decides
    Entity "1" o-- "0..*" Document : documented-by
    Entity "1" o-- "0..*" Runbook : procedure-for

Design notes carried from rev 2 (validated): VMs and LXCs share storage pools via MountVolume on abstract ComputeEntity; not every machine is a Proxmox host (fleet today: 2 PVE hosts, 2 workstations, 1 VPS, 2 VMs, 19 LXCs); Docker containers are first-class (the OS models itself); services attach to any compute entity; documents/runbooks attach to any entity via the root abstract type.

Lifecycles

All lifecycles have explicit terminal states and recovery paths (SA4). Transition preconditions are named checks implemented in Go and referenced by ID from lifecycle_defs.transitions (e.g. no-inbound-edges, backup-verified) — the DB stores which checks gate a transition; the code implements them.

Infrastructure:

stateDiagram-v2
    [*] --> planned : operator creates entity
    planned --> provisioning : IP reserved, storage chosen, doc stub
    planned --> destroyed : cancelled
    provisioning --> active : mesh joined, health answering, doc complete
    provisioning --> failed : provision failed
    active --> migrating : preflight + backup verified
    migrating --> active : post-verify (ingress, mounts checked)
    migrating --> failed : migration failed
    failed --> active : recovered
    failed --> deprecated : written off
    active --> deprecated : replacement live or role retired
    deprecated --> active : un-deprecate (replacement failed)
    deprecated --> destroyed : backups verified, secrets revoked,\ningress removed, zero inbound edges
    destroyed --> [*] : archaeology entry recorded

Signal:

stateDiagram-v2
    [*] --> raised : check fails / agent finding
    raised --> acknowledged : agent or operator sees it
    raised --> muted : operator suppresses (TTL)
    raised --> resolved : condition cleared (auto-resolve)
    acknowledged --> acting : actuator starts execution
    acknowledged --> resolved : manual resolve
    acknowledged --> muted
    acting --> resolved : action + verification passed
    acting --> raised : action failed, re-escalated (retry budget left)
    acting --> failed : permanent failure — needs operator
    failed --> acknowledged : operator retries
    muted --> raised : TTL expired and condition persists
    resolved --> [*]

Execution:

stateDiagram-v2
    [*] --> proposed : classification produced an action
    proposed --> approved : operator approves (if gated)
    proposed --> auto_approved : risk class allows auto-act
    proposed --> denied : operator denies
    approved --> expired : approval TTL ran out
    approved --> executing
    auto_approved --> executing
    executing --> verified : verification passed
    executing --> failed : execution or verification failed
    executing --> timed_out
    executing --> cancelled : operator abort
    timed_out --> verifying : check if command completed anyway
    verifying --> verified
    verifying --> failed
    failed --> rolled_back : rollback procedure executed
    failed --> rollback_failed : rollback also failed — page operator
    verified --> [*] : feedback recorded
    failed --> [*] : feedback recorded
    rolled_back --> [*] : feedback recorded
    rollback_failed --> [*] : feedback recorded
    cancelled --> [*]
    denied --> [*]
    expired --> [*]

Approval: pending → approved | denied | expired; approved → revoked (operator changes mind before execution starts).

Pattern: hypothesized → validated → active → deprecated; hypothesized → invalidated (disproven, terminal); active → invalidated (new evidence contradicts). validated → active requires operator approval (S4).

Skill: drafted → tested → active; active → refined → active (new version); tested → failed → drafted; drafted|active → deprecated.

Cognition loop (observe → decide → act → learn)

flowchart TB
    subgraph observe["Observe"]
        SIG["Signal raised\n(check failed, drift, agent finding)"]
    end
    subgraph decide["Decide"]
        CLASS["Classification\nrisk × blast radius × confidence\n+ recommended action"]
        SKILL_LOOKUP["Skill lookup\nbest-known procedure for\n(entity type, action)"]
        CLASS --> SKILL_LOOKUP
    end
    subgraph act["Act"]
        EXEC["Execution\nrun skill procedure\n(or escalate if none)"]
        VERIFY["Verification"]
        EXEC --> VERIFY
    end
    subgraph learn["Learn"]
        OUTCOME["Outcome evaluation"]
        FEEDBACK["Feedback record"]
        PATTERN["Pattern extraction"]
        SKILL_REFINE["Skill refinement\n(operator gates activation)"]
        OUTCOME --> FEEDBACK --> PATTERN --> SKILL_REFINE
    end
    SIG --> CLASS
    SKILL_LOOKUP --> EXEC
    VERIFY --> OUTCOME
    SKILL_REFINE -. informs next decision .-> SKILL_LOOKUP
    CLASS -- needs approval --> APPROVAL["Approval request → Matrix ✅/❌"]
    APPROVAL -- approved --> EXEC
    APPROVAL -- denied --> RESOLVE["Resolve signal (denied)"]

recommended_action lives on the classification, not the signal (SA6): checks raise facts; the classifier decides what to do about them.

Target architecture

Container stack on mac-mini

One image (oikos:<git-sha>), three long-running roles + init jobs, three Docker networks as trust boundaries (S10):

  • net-front — Caddy-facing: api only.
  • net-data — Postgres + everything that needs it.
  • net-ops — SSH egress: actuator role only (Hermes and API have no SSH).
graph TB
    subgraph mac-mini["mac-mini — Docker host (OrbStack), always-on"]
        subgraph stack["Docker Compose — image oikos:&lt;sha&gt;"]
            MIGRATE["init: oikos migrate + seed\n(one-shot, DDL user)"]
            PG["PostgreSQL 16 + TimescaleDB\nentities/relationships • signals\nexecutions • patterns/skills\npolicy • ontology • hypertables"]
            INF["Infisical (secrets)"]
            API["oikos api\nREST (OpenAPI) + MCP (streamable HTTP)\npolicy enforcement • audit • events\nSSE stream"]
            SCHED["oikos scheduler\nchecks (from check_defs) → signals\n+ actuator: classify → execute/escalate\n+ learning engine\nSSH (restricted key) → fleet"]
            NOTIF["oikos notifier\nMatrix alerts + approval reactions\nDB rendezvous (no RPC)"]
            HERMES["Hermes agent (gateway :8092)\nMCP client → api\nNO SSH keys"]
        end
        DEPLOY["deploy webhook listener\n(HMAC-verified, non-root)"]
    end
    subgraph external["External"]
        CADDY["Caddy (LXC 121)"]
        GITEA["Gitea (LXC 104) + CI"]
        MATRIX["Matrix (LXC 118)"]
        FLEET["hubris / strong / LXCs / VMs"]
        WS["Workstations (Hermes remote)"]
        WATCHDOG["Watchdog cron on apps/105\ncurl /healthz → Matrix ping"]
    end
    MIGRATE --> PG
    API --> PG
    SCHED --> PG
    NOTIF --> PG
    HERMES -- "MCP (bearer token)" --> API
    SCHED -- "SSH (restricted key)" --> FLEET
    NOTIF --- MATRIX
    CADDY -- "reverse_proxy + forward-auth" --> API
    GITEA -- "webhook (HMAC, CI-gated)" --> DEPLOY
    WS -- "mesh, mTLS/token" --> HERMES
    WATCHDOG -. "external heartbeat" .-> API

Container hardening: read-only root filesystems, non-root users, restart: always, stop_grace_period: 30s, images pinned by digest, gcr.io/distroless/static runtime base, CGO_ENABLED=0, pgx (pure Go).

macOS host realities (R3-13)

Docker on macOS runs in a lightweight VM (use OrbStack: fast, auto-starts at login, stable networking). Consequences the deploy must respect:

  • No network_mode: host. All inbound reachability is via published ports.
  • Outbound (actuator SSH → hubris/strong/LXCs): direct over the LAN. Containers reach the LAN via Docker NAT with no special config; SSH provides its own encryption, so there is no reason to route this hop through the mesh.
  • Inbound: mesh-primary, LAN break-glass. NetBird runs on the macOS host.
    • Primary: Caddy and workstations reach the API and Hermes gateway on ports published on the mesh IP (<mesh-ip>:8090 API, :8092 Hermes gateway) — WireGuard-encrypted Caddy→backend hop, mesh membership as a network-level filter, stable addressing, works for roaming workstations.
    • Break-glass: the API port is also published on the mac-mini's LAN IP (<lan-ip>:8090; give mac-mini a DHCP reservation). Safe because the API authenticates every request itself (OIDC JWT / bearer token — S6); the network you arrive from is defense-in-depth, not the auth. This keeps the control plane reachable from inside the house if NetBird's management plane is down after a reboot, and lets the watchdog test the API independently of the mesh. The Hermes gateway stays mesh-only (no LAN binding — it has the weakest application-layer auth story and no break-glass need).
    • Never bind published ports to 0.0.0.0; enumerate mesh IP, LAN IP (API only), and localhost explicitly.
  • Host prep: disable sleep (sudo pmset -a sleep 0 displaysleep 10), auto-restart after power failure (sudo pmset -a autorestart 1), auto-login enabled so OrbStack starts, macOS auto-updates deferred/scheduled (O7).
  • Volumes: keep Postgres data on a named Docker volume (VM-native filesystem), not a bind mount — bind mounts cross the VM boundary and are slow.

Multi-node path (designed for, not built)

All roles are stateless; Postgres is the only stateful service. Scaling later means: run oikos scheduler on another node pointed at the same DB (checks can be partitioned by a zone attribute on check_defs); run oikos api behind Caddy on N nodes. SELECT … FOR UPDATE SKIP LOCKED + advisory locks already make the work queue multi-consumer-safe. Postgres remains the accepted SPOF (mitigated by backup/DR, below), with an upgrade path to streaming replication if ever needed.

API contract

Contract-first (R3-2)

  • api/openapi.yaml (OpenAPI 3.1) is the source of truth. CI fails if handlers drift from the spec.
  • Server: oapi-codegen strict-server stubs on chi + stdlib net/http.
  • Clients: the homelab CLI and future web UIs consume generated clients (Go / TypeScript via openapi-typescript). The spec is served at GET /api/v1/openapi.yaml and human docs at GET /api/v1/docs (Redoc/Scalar static page) — a UI developer needs nothing but the running API.

Conventions (R3-3)

  • Versioning: everything under /api/v1. Additive-only within v1 (new fields, new endpoints); breaking changes ship as /api/v2 side by side with a deprecation window.
  • Errors: RFC 9457 application/problem+json: {"type":"https://oikos.dev/errors/invalid-transition","title":"invalid lifecycle transition","status":409,"detail":"...","instance":"/api/v1/entities/…","errors":[{field,reason}]}. Domain sentinel errors map centrally: ErrNotFound→404, ErrInvalidTransition→409, ErrApprovalRequired→403, ErrAutonomyBlocked→403, ErrConflict→409, ErrCircuitOpen→503, validation → 422.
  • Lists: uniform envelope {"items":[…],"next_cursor":"…"}. Cursor pagination everywhere (?cursor=&limit=, default 50, max 200): keyset on (created_at,id) for entity-ish tables, on ts for hypertables. MCP tools take the same limit.
  • Idempotency: unsafe POSTs (/executions, /approvals/{id}/decision) accept an Idempotency-Key header; keys + response snapshots stored 24h; replay returns the original response. Agents retry safely (R3-3).
  • Optimistic concurrency: mutable resources carry a version; GET returns ETag; PATCH/PUT require If-Match, mismatch → 412. UIs can safely edit.
  • Timestamps: TIMESTAMPTZ in DB, RFC 3339 UTC on the wire.
  • CORS: config-driven origin allowlist (empty by default; future UI origins added via config, not code).
  • Rate limiting (M2): token bucket per authenticated actor; tighter budget on /executions; per-(entity,action) cooldown enforced in the actuator besides.

AuthN/AuthZ

Caller Mechanism Scope
Operator (browser/CLI) Authentik OIDC; Caddy forward-auth and JWT validated in API middleware (defense in depth, S6/SA10) role operator (full) or viewer (read-only)
Hermes (MCP) Static bearer token from Infisical, dedicated Docker network, HMAC on requests (S2) role agent — read tools + POST /executions (which is always policy-gated)
Internal roles (scheduler/notifier) Direct DB with least-privilege DB users; no API hop n/a
/healthz, /metrics No auth, no audit; not exposed via Caddy (SG18) n/a

Roles are claims checked per-route in generated middleware; the spec annotates each operation with its required scope, so future UIs can render capability-aware.

REST surface (summary; the OpenAPI file is normative)

Inventory + ontology:

  • GET/POST /api/v1/entities, GET/PATCH /api/v1/entities/{id-or-slug} (PATCH covers attribute edits and lifecycle transitions; transition legality validated against lifecycle_defs)
  • GET /api/v1/entities/{id}/relations, GET /api/v1/graph?root=&depth=&rel_type={nodes:[…],edges:[…]} for UI visualization (R3-3)
  • GET /api/v1/ontology (types, relationship types, lifecycles), POST /api/v1/ontology/entity-types, PATCH /api/v1/ontology/entity-types/{name} (policy-gated config_mutation; deprecate-not-delete while instances exist, D3)

Operations:

  • GET /api/v1/signals, POST /api/v1/signals/{id}/ack|resolve|mute
  • GET /api/v1/checks, POST /api/v1/checks, PATCH /api/v1/checks/{id} (R3-7)
  • GET /api/v1/approvals, POST /api/v1/approvals/{id}/decision
  • POST /api/v1/executions (classify → approval check → enqueue; SG15), GET /api/v1/executions/{id}, POST /api/v1/executions/{id}/cancel
  • GET /api/v1/classifications?signal_id=…

Learning (operator safety valves, SG7):

  • GET /api/v1/patterns, PATCH /api/v1/patterns/{id} (activate/invalidate — policy-gated config_mutation, audit-logged)
  • GET /api/v1/skills, GET /api/v1/skills/{id}/versions, PATCH /api/v1/skills/{id}

Policy (dual-control, S3):

  • GET /api/v1/policy/risk-classes|approval-rules|autonomy
  • PATCH on any policy resource creates a meta-approval; the change applies only after operator approval; before/after hash audit-logged.

Knowledge + observability:

  • GET /api/v1/knowledge/search?q=, GET /api/v1/knowledge/{entity_id}
  • GET /api/v1/metrics?entity_id=&metric=&from=&to=&rollup=raw|1h|1d (auto-selects resolution by range), GET /api/v1/trends/{entity_id}
  • GET /api/v1/audit?actor=&entity_id=&action=&correlation_id=&from=&to=
  • GET /api/v1/events?type=&entity_id=&severity=&from=&to= and GET /api/v1/events/stream (SSE; Last-Event-ID resume; heartbeat comments; bounded per-subscriber buffers, drop-oldest, P6)
  • GET /api/v1/agent-activity, GET /api/v1/health (fleet summary + trends)
  • GET /api/v1/export — regenerates the three seed YAMLs from DB (round-trip tested, D6)

MCP interface (R3-10)

  • Official github.com/modelcontextprotocol/go-sdk, Streamable HTTP transport, mounted at /mcp on the same binary, bearer-token auth, dedicated network.
  • Tools delegate to the identical service layer as REST (one behavior, two protocols): get_entity, list_entities, get_relations, get_blast_radius, search_knowledge, get_signal_history, get_patterns, get_skills, request_execution, query_metrics, get_trend, get_audit_trail, get_event_timeline, get_agent_activity, get_health_summary.
  • Docs/runbooks additionally exposed as MCP resources (URI = entity slug) so Hermes can attach them as context without a tool round-trip.

DB-native configuration

seeds/ontology.yaml, seeds/inventory.yaml, seeds/policy.yaml bootstrap the DB and serve DR; afterwards the DB is authoritative and editable via API. Ingest runs in the one-shot init container (A4): each seed file applies in a single transaction; a seed_versions table records applied (file, content-hash) so unchanged seeds are skipped; GET /api/v1/export regenerates the YAMLs for commit. Round-trip (seed → DB → export → DB) must be byte-stable — tested in CI.

Config hierarchy (A5): compiled defaults → config file → env vars → Infisical (secrets only — never plain config). Each role's required keys documented in docs/operations/config.md.

Repo layout

/
├── docker-compose.yml
├── Makefile                    # build, test, lint, generate, deploy targets
├── go.mod / go.sum
├── sqlc.yaml
├── .golangci.yml
├── api/
│   └── openapi.yaml            # THE API contract (R3-2)
├── cmd/
│   └── oikos/                  # single binary: api | scheduler | notifier | all |
│       └── main.go             #   migrate | seed | export  (R3-4)
├── internal/
│   ├── domain/                 # pure domain types + state machines + sentinel errors
│   │   ├── entity.go signal.go execution.go classification.go
│   │   ├── pattern.go skill.go approval.go check.go errors.go
│   ├── db/                     # pgx pool, sqlc output, repositories (models never escape)
│   │   └── queries/            # sqlc SQL
│   ├── ontology/               # type hierarchy, validation, graph traversal, seed ingest/export
│   ├── httpapi/                # oapi-codegen server impl, middleware (auth, audit,
│   │                           #   idempotency, rate-limit, problem+json mapping), SSE
│   ├── mcp/                    # MCP server (official SDK) over the same services
│   ├── service/                # shared service layer used by httpapi + mcp + loops
│   ├── policy/                 # classify, approve (tokens), autonomy, meta-approval
│   ├── scheduler/              # check runner (check_defs → signals/metrics), dedup, flap
│   ├── actuator/               # queue consumer, SSH exec, verify, circuit breaker, locks
│   ├── learning/               # feedback, pattern extraction, skill refinement
│   ├── notifier/               # interface + matrix impl (DB rendezvous)
│   ├── observability/          # slog setup, metrics, audit, events, correlation
│   └── config/
├── migrations/                 # golang-migrate, embedded, forward-only (D5/O1)
├── seeds/                      # ontology.yaml, inventory.yaml, policy.yaml
├── docs/
│   ├── adr/                    # MADR records (R3-16)
│   ├── operations/             # backup-restore.md, dr.md, config.md, runbooks
│   └── …                       # narrative docs (ingested into knowledge graph)
├── hermes/                     # config.yaml, SOUL.md, skills/homelab-ops/SKILL.md
├── compose/
│   ├── oikos/Dockerfile        # one multi-stage Dockerfile for the binary
│   └── postgres/               # timescale/timescaledb:2-pg16 config
└── scripts/                    # migrate-sops.sh, import-legacy.sh, deploy.sh, watchdog.sh

Developer experience: make generate (oapi-codegen + sqlc), make test (unit + testcontainers), make dev (docker compose --profile dev up with seeded fake data + oikos all), .env.example committed.

Database schema

Consolidated migrations — all rev-2 fixes applied inline. Forward-only (no down.sql; compensating migrations for rollback, plus pre-deploy dumps). Runner: golang-migrate via oikos migrate in the init container with a DDL-only DB user; runtime roles get DML-only users (SA9). updated_at maintained by a shared trigger.

001 — Ontology meta-schema (with inheritance, R3-1)

CREATE TABLE lifecycle_defs (
    id            TEXT PRIMARY KEY,          -- 'infrastructure', 'signal', ...
    states        TEXT[] NOT NULL,
    default_state TEXT NOT NULL,
    terminal_states TEXT[] NOT NULL DEFAULT '{}',
    transitions   JSONB NOT NULL,            -- {"from":{"to":{"requires":["no-inbound-edges",...]}}}
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE entity_types (
    name          TEXT PRIMARY KEY,          -- 'compute-entity', 'machine', 'proxmox-host'
    parent_type   TEXT REFERENCES entity_types(name),   -- is-a hierarchy (R3-1)
    is_abstract   BOOLEAN NOT NULL DEFAULT false,       -- abstract types can't be instantiated
    domain        TEXT NOT NULL,             -- 'physical','compute','network','storage',
                                             -- 'software','identity','policy','cognition'
    layer         TEXT NOT NULL CHECK (layer IN ('meta','infrastructure','governance','cognition')),
                                             -- 'meta' is reserved for the abstract root type 'entity'
    description   TEXT,
    lifecycle_id  TEXT REFERENCES lifecycle_defs(id),
    attribute_schema JSONB,                  -- JSON Schema for entities.attributes (D2)
    schema_version INTEGER NOT NULL DEFAULT 1,
    status        TEXT NOT NULL DEFAULT 'active',  -- 'active'|'deprecated'; no hard delete
                                                   -- while instances exist (D3)
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE relationship_types (
    name          TEXT PRIMARY KEY,          -- 'hosts', 'provides', 'depends-on'
    inverse       TEXT,
    source_type   TEXT NOT NULL REFERENCES entity_types(name),  -- MAY be abstract;
    target_type   TEXT NOT NULL REFERENCES entity_types(name),  -- validation walks hierarchy
    cardinality   TEXT NOT NULL CHECK (cardinality IN          -- source→target multiplicity
                    ('one-to-one','one-to-many','many-to-one','many-to-many')),
    description   TEXT,
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE seed_versions (               -- A4: skip unchanged seed files
    file          TEXT PRIMARY KEY,
    content_hash  TEXT NOT NULL,
    applied_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

Validation semantics (app layer, internal/ontology):

  • Instantiating an is_abstract type is rejected.
  • A relationship (s, t, type) is valid iff type_of(s) is source_type or a descendant of it (same for target). The hierarchy is small; resolved in-memory with a cached type tree.
  • Cardinality enforced by partial unique indexes where expressible (one-to-manyUNIQUE (target_id, type); one-to-one → unique on both ends) plus app-layer checks for the rest.

002 — Entity instances (UUIDv7 + slug, R3-5/D1)

CREATE TABLE entities (
    id            UUID PRIMARY KEY,          -- UUIDv7 generated in Go (time-ordered)
    slug          TEXT NOT NULL UNIQUE,      -- 'host:hubris', 'service:caddy' — human/API handle
    type          TEXT NOT NULL REFERENCES entity_types(name),
    name          TEXT NOT NULL,
    state         TEXT,                      -- lifecycle state
    attributes    JSONB NOT NULL DEFAULT '{}',   -- validated against attribute_schema
    maintenance_until TIMESTAMPTZ,           -- R3-8: suppress signals + auto-act while set
    version       INTEGER NOT NULL DEFAULT 1,    -- optimistic lock / ETag source
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    UNIQUE (type, name)
);
CREATE INDEX idx_entities_type  ON entities(type);
CREATE INDEX idx_entities_state ON entities(state);
CREATE INDEX idx_entities_attrs ON entities USING GIN(attributes);

CREATE TABLE relationships (
    source_id     UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT,
    target_id     UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT,
    type          TEXT NOT NULL REFERENCES relationship_types(name),
    attributes    JSONB,
    valid_from    TIMESTAMPTZ NOT NULL DEFAULT now(),  -- D7: temporal edges
    valid_to      TIMESTAMPTZ,                          -- NULL = current
    PRIMARY KEY (source_id, target_id, type, valid_from)
);
CREATE INDEX idx_rel_source ON relationships(source_id) WHERE valid_to IS NULL;
CREATE INDEX idx_rel_target ON relationships(target_id) WHERE valid_to IS NULL;
CREATE INDEX idx_rel_type   ON relationships(type)      WHERE valid_to IS NULL;

-- Cycle-safe traversal (P1): path accumulator prevents revisits; depth capped.
CREATE OR REPLACE FUNCTION blast_radius(start_id UUID, max_depth INT DEFAULT 3,
                                        rel_types TEXT[] DEFAULT NULL)
RETURNS TABLE(entity_id UUID, depth INT) AS $$
    WITH RECURSIVE walk AS (
        SELECT start_id AS entity_id, 0 AS depth, ARRAY[start_id] AS path
        UNION ALL
        SELECT r.target_id, w.depth + 1, w.path || r.target_id
        FROM relationships r
        JOIN walk w ON r.source_id = w.entity_id
        WHERE w.depth < LEAST(max_depth, 5)
          AND r.valid_to IS NULL
          AND NOT r.target_id = ANY(w.path)
          AND (rel_types IS NULL OR r.type = ANY(rel_types))
    )
    SELECT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
$$ LANGUAGE sql STABLE;

Entities are never hard-deleted while edges exist (ON DELETE RESTRICT); decommission is the lifecycle path (… → destroyed), and destroyed entities remain as archaeology.

Dual-entity pattern (SA1): every cognition object (signal, classification, execution, feedback, pattern, skill, approval, check) gets an entities row (so the graph is traversable: triggers, produces, contributes-to edges live in relationships) and a typed table below whose PK references entities(id) for indexed querying.

003 — Operations (signals, checks, approvals, status)

CREATE TABLE check_defs (                    -- R3-7: probes as data
    entity_id     UUID PRIMARY KEY REFERENCES entities(id),
    target_id     UUID REFERENCES entities(id),        -- what it checks (NULL + target_type = type-scoped)
    target_type   TEXT REFERENCES entity_types(name),
    kind          TEXT NOT NULL,             -- 'http','tcp','disk','cert-expiry','drift','ssh-script'
    config        JSONB NOT NULL DEFAULT '{}',  -- validated per-kind JSON Schema
    interval_s    INTEGER NOT NULL DEFAULT 600,
    timeout_s     INTEGER NOT NULL DEFAULT 10,
    zone          TEXT,                      -- multi-node partitioning later
    enabled       BOOLEAN NOT NULL DEFAULT true,
    updated_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE signals (
    entity_id         UUID PRIMARY KEY REFERENCES entities(id),
    kind              TEXT NOT NULL,
    severity          TEXT NOT NULL CHECK (severity IN ('info','warning','critical')),
    target_entity_id  UUID REFERENCES entities(id),
    check_id          UUID REFERENCES check_defs(entity_id),
    evidence          TEXT,
    likely_cause      TEXT,
    state             TEXT NOT NULL DEFAULT 'raised',
    occurrence_count  INTEGER NOT NULL DEFAULT 1,      -- R3-8: dedup counting
    first_seen_at     TIMESTAMPTZ NOT NULL DEFAULT now(),
    last_seen_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    flap_count        INTEGER NOT NULL DEFAULT 0,      -- resolve→re-raise cycles
    hold_down_until   TIMESTAMPTZ,                     -- flap suppression window
    mute_until        TIMESTAMPTZ,
    created_at        TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at        TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- R3-8: at most ONE open signal per (target, kind) — repeats update the open row
CREATE UNIQUE INDEX uq_signals_open ON signals(target_entity_id, kind)
    WHERE state NOT IN ('resolved','failed');
CREATE INDEX idx_signals_state ON signals(state);

CREATE TABLE approvals (
    entity_id         UUID PRIMARY KEY REFERENCES entities(id),
    subject_entity_id UUID REFERENCES entities(id),    -- entity to act on
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    kind              TEXT NOT NULL DEFAULT 'execution',  -- 'execution'|'policy-change'|'pattern-activation'
    payload           JSONB,                            -- e.g. the proposed policy diff
    status            TEXT NOT NULL DEFAULT 'pending',
    token_hash        TEXT,                             -- S5: single-use HMAC token, stored hashed
    expires_at        TIMESTAMPTZ NOT NULL,
    decided_at        TIMESTAMPTZ,
    decided_by        UUID REFERENCES entities(id),     -- person entity
    created_at        TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE entity_status (                 -- R3-6: replaces state_snapshots (P5)
    entity_id     UUID PRIMARY KEY REFERENCES entities(id),
    health        TEXT NOT NULL DEFAULT 'unknown',  -- healthy|degraded|down|unknown
    last_check_at TIMESTAMPTZ,
    details       JSONB NOT NULL DEFAULT '{}',
    updated_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- health HISTORY is the 'health' metric in metric_samples (retained + rolled up)

CREATE TABLE idempotency_keys (              -- R3-3
    key           TEXT NOT NULL,
    actor         TEXT NOT NULL,
    request_hash  TEXT NOT NULL,
    response_code INTEGER,
    response_body JSONB,
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    PRIMARY KEY (actor, key)
);
-- pruned by the scheduler after 24h

Approval tokens (S5): token = HMAC(approval_id ‖ subject ‖ action ‖ risk_class ‖ nonce, secret), single-use, stored hashed, TTL-bound; Matrix carries only the approval ID + decision; verification is server-side.

004 — Cognition (classifications, executions, learning)

CREATE TABLE classifications (               -- SA5: every autonomous decision persisted
    entity_id         UUID PRIMARY KEY REFERENCES entities(id),
    signal_entity_id  UUID REFERENCES signals(entity_id),
    target_entity_id  UUID REFERENCES entities(id),
    action            TEXT NOT NULL,
    recommended_action JSONB,                -- SA6: lives here, not on the signal
    risk_class        TEXT NOT NULL,
    route             TEXT NOT NULL CHECK (route IN ('auto-act','escalate','hold')),
    blast_radius      UUID[],
    pattern_confidence REAL,
    skill_id          UUID,                  -- skill entity matched (if any)
    autonomy_check    TEXT,                  -- 'allowed' | 'blocked: <reason>'
    reasoning         JSONB NOT NULL,
    correlation_id    TEXT NOT NULL,
    created_at        TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE executions (
    entity_id         UUID PRIMARY KEY REFERENCES entities(id),
    classification_id UUID REFERENCES classifications(entity_id),
    signal_entity_id  UUID REFERENCES signals(entity_id),
    target_entity_id  UUID REFERENCES entities(id),
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL,
    approval_id       UUID REFERENCES approvals(entity_id),
    agent_id          UUID REFERENCES entities(id),
    skill_id          UUID,                  -- + version pinned at execution time (SG9)
    skill_version     INTEGER,
    status            TEXT NOT NULL DEFAULT 'proposed',
    result            JSONB,
    duration_ms       INTEGER,
    verified          BOOLEAN NOT NULL DEFAULT false,
    correlation_id    TEXT NOT NULL,
    started_at        TIMESTAMPTZ,
    completed_at      TIMESTAMPTZ,
    created_at        TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_exec_target ON executions(target_entity_id);
CREATE INDEX idx_exec_status ON executions(status);

CREATE TABLE feedback (
    entity_id     UUID PRIMARY KEY REFERENCES entities(id),
    execution_id  UUID NOT NULL REFERENCES executions(entity_id),
    outcome       TEXT NOT NULL CHECK (outcome IN ('success','failure','partial','unexpected')),
    observation   TEXT,
    lesson        TEXT,
    unexpected_side_effects TEXT[],
    tags          TEXT[],
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_feedback_ts ON feedback(created_at);   -- P4: watermark scans

CREATE TABLE patterns (
    entity_id     UUID PRIMARY KEY REFERENCES entities(id),
    applies_type  TEXT NOT NULL REFERENCES entity_types(name),
    action        TEXT NOT NULL,
    pattern       TEXT NOT NULL,
    confidence    REAL NOT NULL DEFAULT 0,   -- Wilson lower bound, capped by sample size (S4)
    evidence_count INTEGER NOT NULL DEFAULT 0,
    success_count INTEGER NOT NULL DEFAULT 0,
    failure_count INTEGER NOT NULL DEFAULT 0,
    status        TEXT NOT NULL DEFAULT 'hypothesized',
    quarantined   BOOLEAN NOT NULL DEFAULT false,   -- S4: anomalous feedback bursts
    version       INTEGER NOT NULL DEFAULT 1,       -- D4: optimistic lock
    last_validated_at TIMESTAMPTZ,
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    UNIQUE (applies_type, action)
);
-- Counters updated atomically: UPDATE … SET evidence_count = evidence_count + 1 (D4)

CREATE TABLE skills (
    entity_id     UUID NOT NULL REFERENCES entities(id),
    version       INTEGER NOT NULL DEFAULT 1,       -- SG9: history preserved
    name          TEXT NOT NULL,
    procedure     JSONB NOT NULL,           -- R3-9: structured, schema-validated (below)
    applies_type  TEXT REFERENCES entity_types(name),
    action        TEXT NOT NULL,
    pattern_ids   UUID[],
    status        TEXT NOT NULL DEFAULT 'drafted',
    success_rate  REAL,
    changed_by    UUID,                     -- agent or person entity
    change_reason TEXT,
    last_used_at  TIMESTAMPTZ,
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    PRIMARY KEY (entity_id, version)
);

Skill procedure format (R3-9) — validated by JSON Schema at write time; the actuator executes it deterministically; markdown for humans is rendered from it:

{
  "params_schema": { "type": "object", "properties": { "unit": {"type":"string"} }, "required": ["unit"] },
  "steps": [
    { "name": "restart unit",
      "runner": "ssh", "target": "{{ .host }}",
      "command": "systemctl restart {{ .unit }}",
      "timeout_s": 60 }
  ],
  "verify": [
    { "runner": "ssh", "target": "{{ .host }}",
      "command": "systemctl is-active {{ .unit }}",
      "expect": { "exit_code": 0, "stdout_contains": "active" },
      "retry": { "attempts": 3, "delay_s": 10 } }
  ],
  "rollback": [],
  "expected_duration_s": 30,
  "known_failure_modes": ["unit masked", "dependency service down"]
}

Templates are Go text/template over validated params; the actuator refuses any command not generated from a stored skill/runbook step (no free-form agent shell).

005 — Policy

CREATE TABLE risk_classes (
    name              TEXT PRIMARY KEY,      -- read_only, reversible_low, config_mutation, destructive
    description       TEXT,
    approval_required TEXT NOT NULL DEFAULT 'none',   -- none|operator|operator_confirmed
    autonomy_allowed  BOOLEAN NOT NULL DEFAULT false
);

CREATE TABLE approval_rules (
    id                UUID PRIMARY KEY,
    entity_type       TEXT REFERENCES entity_types(name),   -- may be abstract (R3-1)
    action            TEXT NOT NULL,
    risk_class        TEXT NOT NULL REFERENCES risk_classes(name),
    autonomy_level    TEXT NOT NULL DEFAULT 'escalate' CHECK
                        (autonomy_level IN ('auto','escalate','never')),
    scope_entity      UUID REFERENCES entities(id),          -- optional per-entity override
    version           INTEGER NOT NULL DEFAULT 1,
    updated_at        TIMESTAMPTZ NOT NULL DEFAULT now(),
    UNIQUE (entity_type, action, scope_entity)
);

CREATE TABLE autonomy_settings (
    key               TEXT PRIMARY KEY,      -- 'global.auto_act', 'never_auto_act.<slug>'
    value             TEXT NOT NULL,
    version           INTEGER NOT NULL DEFAULT 1,
    updated_at        TIMESTAMPTZ NOT NULL DEFAULT now()
);

Rule resolution: most-specific wins (scope_entity > concrete type > ancestor type via the hierarchy). Policy mutations are dual-controlled (S3): the API writes a policy-change approval; on operator approval the change applies in one transaction with a before/after hash written to audit_log; at startup each role verifies the policy hash against last-known-good and raises a critical signal on mismatch.

006 — Observability (TimescaleDB)

Hypertable PKs include the time column (SG1); TimescaleDB DDL is idempotent (if_not_exists => TRUE, exception-guarded policies — SG3); CAGGs avoid array_agg (SG2).

CREATE EXTENSION IF NOT EXISTS timescaledb;

CREATE TABLE metric_samples (
    ts          TIMESTAMPTZ NOT NULL,
    entity_id   UUID NOT NULL,
    metric      TEXT NOT NULL,               -- 'health','disk_usage_pct','probe_latency_ms',
                                             -- 'api_latency_ms','pattern_confidence',…
    value       DOUBLE PRECISION NOT NULL,
    tags        JSONB NOT NULL DEFAULT '{}'
);
SELECT create_hypertable('metric_samples','ts',
    chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
CREATE INDEX idx_metrics_entity_ts ON metric_samples(entity_id, ts DESC);
CREATE INDEX idx_metrics_metric_ts ON metric_samples(metric, ts DESC);
DO $$ BEGIN
    PERFORM add_retention_policy('metric_samples', INTERVAL '90 days');
EXCEPTION WHEN OTHERS THEN NULL; END $$;

CREATE MATERIALIZED VIEW metric_rollups_1h WITH (timescaledb.continuous) AS
    SELECT time_bucket('1 hour', ts) AS bucket, entity_id, metric,
           avg(value) AS avg_value, min(value) AS min_value,
           max(value) AS max_value, count(*) AS sample_count
    FROM metric_samples GROUP BY bucket, entity_id, metric;
-- + metric_rollups_1d identically; refresh policies 1h/1d; rollups kept 1 year

CREATE TABLE audit_log (
    id          BIGINT GENERATED ALWAYS AS IDENTITY,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    actor_type  TEXT NOT NULL,               -- agent|operator|system|scheduler
    actor_id    UUID,                        -- resolves to a real entity (SA3)
    action      TEXT NOT NULL,
    entity_id   UUID,
    method      TEXT, path TEXT, status_code INTEGER,
    detail      JSONB NOT NULL DEFAULT '{}',
    source_ip   TEXT,
    correlation_id TEXT,
    PRIMARY KEY (id, ts)                     -- SG1
);
SELECT create_hypertable('audit_log','ts',
    chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- indexes on (actor_type,actor_id,ts), (entity_id,ts), (correlation_id); 365d retention

CREATE TABLE events (
    id          BIGINT GENERATED ALWAYS AS IDENTITY,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    type        TEXT NOT NULL,               -- 'signal.raised','execution.completed',…
    entity_id   UUID,
    severity    TEXT NOT NULL DEFAULT 'info',
    source      TEXT NOT NULL,
    data        JSONB NOT NULL DEFAULT '{}',
    correlation_id TEXT,
    PRIMARY KEY (id, ts)
);
SELECT create_hypertable('events','ts',
    chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- 90d retention; NOTIFY trigger fires after commit for the SSE stream (SG10)

CREATE TABLE agent_activity (
    id          BIGINT GENERATED ALWAYS AS IDENTITY,
    ts          TIMESTAMPTZ NOT NULL DEFAULT now(),
    agent_id    UUID NOT NULL,
    session_id  TEXT,
    activity_type TEXT NOT NULL,             -- tool_call|reasoning|decision|mcp_query|escalation
    tool_name   TEXT, entity_id UUID,
    input_summary TEXT, output_summary TEXT, -- truncated 500 chars
    duration_ms INTEGER, token_count INTEGER, success BOOLEAN,
    correlation_id TEXT,
    PRIMARY KEY (id, ts)
);
SELECT create_hypertable('agent_activity','ts',
    chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- 90d retention

-- R3-11: ledger is a VIEW, not a fourth write path
CREATE VIEW ledger AS
SELECT e.created_at AS ts, e.entity_id AS execution_id, e.target_entity_id,
       e.action, e.risk_class, e.status, e.verified,
       c.route, c.reasoning, a.status AS approval_status, a.decided_by,
       e.agent_id, e.correlation_id
FROM executions e
LEFT JOIN classifications c ON c.entity_id = e.classification_id
LEFT JOIN approvals a       ON a.entity_id = e.approval_id;

Retention summary: metrics 90d raw / 1y rollups; audit 1y; events 90d; agent activity 90d; signals/executions/classifications/feedback permanent (they're the learning corpus and small); idempotency_keys 24h (scheduler prune job).

Control loop

Scheduler (Observe)

  • Loads enabled check_defs; each check runs on its own interval with jitter; a bounded worker pool (errgroup.SetLimit) caps concurrency; per-check timeouts (P2).
  • Check kinds implemented in Go, config-driven: http, tcp, disk (SSH df), cert-expiry, drift (DB inventory vs live state — mismatches raise drift signals with the observed diff as evidence), ssh-script (allowlisted).
  • Every run writes metrics (health, probe_latency_ms, …) and updates entity_status in place (R3-6).
  • Signal dedup (R3-8): failure upserts against the partial unique index — an existing open signal gets occurrence_count+1, last_seen_at=now(). Recovery auto-resolves. A resolve→re-raise cycle increments flap_count; after 3 cycles/1h the signal enters hold-down (no notifications, no auto-act) and a flapping meta-signal is raised for the operator.
  • Entities with maintenance_until > now() still get metrics but no signals and are excluded from auto-act.
  • Housekeeping jobs: daily pg_dump + rclone push, idempotency-key prune, feedback watermark advance, secrets-expiry check (M5).

Actuator (Act)

  • Consumes signals with route='auto-act' classifications via SELECT … FOR UPDATE SKIP LOCKED; per-target serialization with pg_advisory_xact_lock(hashtext(target_entity_id::text)) (SG5).
  • Executes only stored skill/runbook procedures (R3-9) over SSH with a restricted key: dedicated keypair, command=/from= constrained in authorized_keys on targets, mounted read-only into the scheduler/actuator container only (S1). Phase 3 formalizes this as the gateway: all execution flows through /executions; Hermes never touches SSH.
  • Context-aware SSH (SG13): session started in a goroutine, ctx.Done() closes session+client to unblock; SSH errors classified (network → retryable/circuit, auth → fatal alert, non-zero exit → failed, timeout → timed_out → verifying).
  • Circuit breaker per target host (M4): N consecutive failures → open circuit, exponential backoff, raise target-unreachable signal instead of piling failed executions into the learning corpus.
  • Loop guard: retry budget per (entity, action) from execution history; rate cooldown per (entity, action).
  • Autonomy kill-switch consulted every pass (global.auto_act, never_auto_act.<slug>).
  • Graceful shutdown (SG4): signal.NotifyContext; stop intake → 30s drain → if an execution is in flight, mark failed ("shutdown interrupted") + emit feedback → close pool. Compose sets stop_grace_period: 30s.

Learning engine

Concrete algorithm (P4 + S4 guardrails):

  1. Hourly, read feedback past the watermark, joined to executions, grouped by (applies_type, action).
  2. Update pattern counters atomically; recompute confidence as the Wilson score lower bound of success rate (conservative for small N), additionally capped by min(confidence, evidence_count/5) so nothing looks confident before 5 samples.
  3. Status: hypothesized (N<5) → validated (N≥5 and confidence ≥ 0.7, emits pattern.validated event + notification) → active only via operator PATCH (policy-gated config_mutation).
  4. Anomaly quarantine: >10 identical-outcome feedback rows within 1h for one (type, action) → quarantined=true, meta-signal for review.
  5. Skill refinement: an active pattern with confidence > 0.7 drafts/refines a skill (new version row; changed_by, change_reason recorded). Skills follow their own lifecycle; activation is operator-gated. No skill ever auto-promotes an action into destructive autonomy — hard-coded, not policy data.
  6. Classifier consumption: pattern confidence and skill existence feed the auto-act/escalate score; frequent-failure patterns lower confidence.

Notifier

Notifier interface (SendAlert / SendApprovalRequest / decision intake); Matrix implementation posts approvals with / reactions and writes decisions directly to the approvals table — the DB is the rendezvous, no service-to-service calls, so pending approvals survive restarts of either side (SA7/A7). If the notifier is down, the operator's alternative path is the REST /approvals endpoint.

Security model

Threat model (documented in docs/adr/0007-threat-model.md):

Boundary Mechanism
Internet/mesh → API Caddy (TLS) + Authentik forward-auth and in-API OIDC JWT validation (Caddy compromise ≠ API compromise)
LAN → API (break-glass) Same in-API auth (OIDC JWT / bearer) — network origin is defense-in-depth, never the auth. Plaintext hop accepted for emergency/watchdog use only; routine traffic uses the mesh
Workstation → Hermes gateway Mesh membership (network) + gateway token/mTLS (application); mesh-only, no LAN binding
Hermes → API (MCP) Dedicated Docker network + static bearer token + HMAC
Roles → Postgres Least-privilege DB users (DDL only in init; DML per role), TLS on the Docker network (S7)
Actuator → fleet Restricted SSH key (command=,from=), actuator container only; full gateway in Phase 3 (S1/S10)
Gitea → deploy HMAC-signed webhook, localhost/mesh-bound listener, non-root deploy user, CI-gated (S8/M1)
Learning → policy Structurally impossible: learning role's DB user has no write grants on policy tables; changes route through approvals (S3/S4)
Approvals Single-use HMAC tokens, hashed at rest, TTL (S5)
Secrets Infisical with machine identities; bootstrap root of trust = master key in mac-mini Keychain, backed up offline; one SOPS age key retained for DR until a restore drill passes (S9)

Rotation cadences (M5): SSH actuator key 6mo, Infisical machine tokens 90d, MCP bearer 90d, webhook HMAC 1y — each with a documented procedure and a scheduler check that raises a signal 2 weeks before expiry. Supply chain (M6): images pinned by digest, govulncheck + golangci-lint in CI, distroless runtime.

Observability

All capture goes to Postgres (hypertables above) through internal/observability:

  • Metrics — check results, API request count/latency, Go runtime stats, learning metrics (pattern confidence, skill success rate, auto-act vs escalation ratio), agent token/tool-call counts. Also exposed as a Prometheus-format /metrics endpoint (internal only) so Grafana/Prometheus can attach later without schema work.
  • Audit — middleware on every mutating REST call + MCP tool call + actuator SSH command; policy changes carry before/after hashes; operator identity from OIDC claims (M3).
  • Events — emitted in the same transaction as the state change (SG10); post-commit NOTIFY feeds the SSE stream; in-process bus covers API-local events (SG8). SSE subscribers get bounded buffers with drop-oldest + heartbeats (P6); delivery is best-effort, history via GET /events.
  • Correlation — a correlation_id is minted at signal creation (or API request) and propagated via context.Context through classification → approval → execution → SSH → verification → feedback, linking audit + events end-to-end.
  • Loggingslog JSON to stdout; every line carries service, correlation_id, entity_id where applicable; debug=true enables probe payloads/SQL/classification reasoning.
  • Agent self-inspection — MCP tools let Hermes query its own history, trends, audit trail, and efficiency (token usage over time).

SLOs (M7): check interval 10min ±1min; API p99 < 200ms; deploy < 5min; alert delivery < 30s; watchdog detection < 5min.

Operations

Backup / restore / DR (A3, O3, O4)

  • Daily pg_dump (custom format, compressed) + WAL archiving for PITR; pushed off-host to Proton Drive via rclone (reuse existing rclone credentials from LXC 132 setup). Retention 30 daily + 12 monthly. Pre-deploy dump before every migration run.
  • Infisical: native backup + secrets exported to one SOPS-age-encrypted file as fallback (the retained age key is the DR escape hatch).
  • Monthly automated restore drill: scratch container, pg_restore, run the export round-trip check, alert on failure.
  • DR targets: RTO 4h / RPO 24h. Cold-start runbook (docs/operations/dr.md): fresh machine → OrbStack + Docker → clone repo → restore Infisical → pg_restoredocker compose up -d → verify /healthz + fleet health.

Watchdog (O2)

Cron on apps/105 (outside the stack): every 5min, curl /healthz on both pathshttp://<lan-ip>:8090/healthz (LAN, tests the API itself) and http://<mesh-ip>:8090/healthz (mesh, tests the path Caddy and workstations use) — plus pg_isready; on any failure, post directly to the Matrix webhook naming which path failed (LAN-down = stack problem; mesh-down-LAN-up = NetBird problem). This is the orthogonal "who watches the watcher" channel.

Deploy, release, rollback (O1, O5, M1, R3-12)

  • CI (Gitea Actions): go vet, golangci-lint run, go test ./... -race -cover (coverage gates: ≥80% internal/policy + internal/learning, ≥60% elsewhere), oapi-codegen/sqlc diff check (generated code committed and clean), govulncheck, docker build.
  • Deploy: webhook (HMAC-verified, gated on green CI) → deploy.sh: git pull → pre-deploy pg_dump → build image oikos:<sha> (multi-stage; no host go build, P7) → oikos migrate init job → docker compose up -d --no-deps app roles (Postgres container never recreated on routine deploys) → healthcheck-gated.
  • Migration compatibility: additive-only per deploy window (new columns nullable); new code tolerates previous schema for one deploy.
  • Rollback: retag compose to the previous SHA (last 5 images kept) → if the migration was the problem, pg_restore the pre-deploy dump → docker compose up -d. Runbook in docs/operations/rollback.md.
  • Health checks (O6): API /healthz (DB ping); scheduler freshness ("last successful check pass < 15min") exposed via entity_status self-row; notifier "last poll < 60s"; Hermes gateway ping — all wired into compose healthcheck + watchdog.

Coexistence + cutover (A6)

During coexistence apps/105 is read-only toward shared state (its scheduler disabled once the new stack's checks are verified). Cutover: verify traffic on the new stack → disable apps/105 services → retarget Caddy + Gitea webhooks → cleanup. Rollback after cutover = re-enable apps/105 (its JSONL state frozen at cutover; reconciliation = re-import from Postgres if ever needed).

Phasing — with acceptance criteria (R3-15)

Each phase ends with explicit "done when" checks an implementing agent can run.

Phase 0 — Ontology + contract (no service code):

  • Finalize entity types (incl. abstract hierarchy, Person/Agent/IdentityProvider, Cluster/ComposeStack), relationship types (with endpoint types + cardinality), lifecycles (all terminal states + named preconditions).
  • Write seeds/*.yaml; write api/openapi.yaml v1 for the full REST surface; write ADRs 00010010 from the Decisions table.
  • Done when: seeds lint against a meta-schema validator script; openapi.yaml passes redocly lint; operator has reviewed the lifecycle diagrams + spec.

Phase 1 — Foundation (DB + skeleton):

  • Go module github.com/dtoro/oikos; cmd/oikos skeleton with role subcommands; domain layer + sentinel errors; migrations 001006; oikos migrate + oikos seed (idempotent, transactional, seed_versions); sqlc + repositories; slog; testcontainers harness; daily backup job + watchdog cron installed.
  • Done when: make test green including: seed ingest twice = no-op; export round-trip byte-stable; blast_radius correct on a cyclic fixture; abstract-type instantiation rejected; relationship endpoint validation honors inheritance; hypertables + CAGGs + retention created idempotently; legacy signals/*.jsonl + ledger/*.jsonl imported by scripts/import-legacy.sh.

Phase 2 — API:

  • oapi-codegen server; auth middleware (OIDC JWT + roles, bearer for agent); problem+json mapping; idempotency; ETag/If-Match; rate limiting; audit middleware; transactional event emitter + SSE stream; knowledge ingestion from docs/ (content-hash skip, P3); MCP server (official SDK) over the shared service layer; pattern/skill/policy endpoints with dual-control approvals; CI pipeline live.
  • Done when: spec-conformance tests pass (schemathesis or generated-client round trip); curl /api/v1/entities?type=service returns the fleet; MCP list_entities returns the same data; a mutating call without If-Match on a stale version → 412; replayed Idempotency-Key returns the cached response; SSE stream shows an entity.created event; audit rows carry the OIDC sub.

Phase 3 — Control loop:

  • Scheduler with check_defs runner, dedup/flap/maintenance logic, metrics + entity_status; actuator with restricted SSH key, advisory locks, circuit breaker, retry budgets, graceful shutdown; learning engine per the algorithm above; approval tokens; notifier (Matrix, DB rendezvous).
  • Done when: killing a probed service raises exactly one signal (repeats increment occurrence_count); a flapping fixture enters hold-down; a service-down signal on a reversible_low target auto-restarts, verifies, and the full correlation chain (signal → classification → execution → SSH audit → feedback) is queryable by one correlation_id; a destructive action produces a Matrix approval whose token is single-use; after 5 successful executions a pattern reaches validated and stays there until operator PATCH; kill-switch global.auto_act=off forces escalation.

Phase 4 — Agent (Hermes):

  • Hermes container (gateway :8092, mesh-published, token/mTLS), homelab skills, MCP wiring, agent-activity logging. No SSH keys in this container.
  • Done when: from a workstation over mesh, Hermes answers "what depends on authentik?" via get_blast_radius, requests an execution that routes through /executions policy gating, and its tool calls appear in agent_activity.

Phase 5 — Secrets (Infisical):

  • Infisical up; SOPS migrated; services on machine identities; rotation checks; SOPS-age DR fallback exported.
  • Done when: no service reads SOPS at runtime; restore drill of Infisical backup passes; rotation runbooks written.

Phase 6 — Deploy + cutover:

  • Full pipeline (CI-gated webhook, SHA images, init migrate, healthcheck rollout); Caddy re-point (mcp.hubris.network, oikos.hubris.network → mac-mini mesh :8090); end-to-end verification below; apps/105 disabled per cutover checklist; first monthly restore drill executed.
  • Done when: all 14 verification checks pass; watchdog alert fires when the API is stopped manually; rollback drill (previous SHA + pg_restore) rehearsed once.

Verification (end to end)

  1. Ontology: SELECT * FROM entity_types shows the 3-layer hierarchy incl. abstract types; lifecycle defs match the diagrams; abstract instantiation is rejected via API (422).
  2. DB: migrations 001006 apply idempotently; seeds ingest; re-ingest is a no-op.
  3. API: REST + MCP return identical data for the fleet; problem+json on errors; pagination envelope everywhere.
  4. Scheduler: check pass writes metrics + entity_status; one open signal per (entity, kind) under repeated failure.
  5. Actuator: classify → auto-act or escalate → execute → verify → feedback, all correlation-linked in audit + events.
  6. Learning: after N=5 similar executions a pattern is validated with a Wilson-bounded confidence; operator PATCH activates it; a skill version appears.
  7. Classifier + autonomy: high-confidence pattern → auto-act on reversible_low; kill-switch off → always escalate; never_auto_act.<slug> honored.
  8. Hermes: remote workstation session; MCP tools work; activity logged.
  9. Secrets: services fetch from Infisical; SOPS retired (except DR key).
  10. Deploy: push → CI → webhook → SHA image → migrate init → rolling restart; deploy.triggered/deploy.completed events emitted.
  11. Knowledge: search_knowledge("caddy") returns docs edged to service:caddy.
  12. Observability: metrics/trends/audit/health endpoints return correct shapes; Grafana can read metric_samples directly.
  13. Correlation tracing: one correlation_id reconstructs signal → classification → execution → SSH → verification → feedback.
  14. Cutover: apps/105 stopped; watchdog still alive; production traffic served solely by the Docker stack.

Risks / trade-offs

  • Go rewrite (~4,400 Python lines replaced) — logic carries over 1:1 (see reuse table); mitigated by the phase gates and the contract-first spec.
  • Postgres SPOF — accepted; mitigated by backups/PITR/DR drills; replication is the future path.
  • Learning cold start — by design: the agent escalates everything until patterns validate and the operator activates them; trust is earned.
  • Single-host mac-mini — watchdog + pmset hardening + documented cold start; multi-node path exists when it matters.
  • Infisical bootstrap — SOPS fallback retained until a restore drill passes.

Reuse map (Python → Go)

Existing Becomes
oikos/decide.py internal/policy/classify.go (+ pattern/skill inputs)
oikos/signal.py internal/scheduler signal lifecycle (DB-backed, dedup added)
oikos/approve.py internal/policy/approve.go (tokens) + internal/notifier/matrix.go
oikos/ledger.py executions + ledger view + audit_log
oikos/policy.py + policy.yaml internal/policy + seeds/policy.yaml → tables
oikos/drift.py drift check kind
oikos/relations.py internal/ontology/graph.go (SQL traversal)
oikos/report.py report endpoints over DB
mcp/server.py internal/mcp (official SDK)
bin/homelab generated OpenAPI Go client
oikos/scheduler.py internal/scheduler (goroutine worker pool, checks-as-data)
(new) internal/learning, internal/observability, internal/httpapi

Out of scope (for now)

  • Web UI (the OpenAPI contract + SSE + /graph endpoint are built for it; the UI itself comes later).
  • Multi-node deployment (designed for; see "Multi-node path").
  • Vector embeddings / semantic search (Postgres FTS now; pgvector is a schema-only addition later).
  • WebSocket stream (SSE covers current needs).
  • LLM-assisted skill extraction (patterns are statistical for now).

Appendix A — audit resolution ledger

Every rev-2 finding and where rev 3 resolves it. (Full finding text in git history, rev 2 of this file.)

Finding Resolution
S1, S10 Security model — restricted SSH key in actuator only; network trust zones; Hermes keyless
S2 MCP bearer token + dedicated network (API contract / Security)
S3 Policy dual-control + hash audit + startup self-check (005 / Security)
S4 Operator-gated pattern activation, Wilson + N/5 cap, quarantine, no destructive auto-promotion (Learning engine)
S5 Single-use HMAC approval tokens, hashed (003)
S6, SA10 In-API OIDC JWT validation; Caddy = explicit trust root (AuthN/AuthZ)
S7 TLS to Postgres; per-role DB users (Security)
S8 HMAC webhook, non-root deploy, CI gate (Deploy)
S9 Infisical bootstrap root of trust + SOPS DR fallback (Security / Phase 5)
P1 Cycle-safe blast_radius with path accumulator + depth cap (002)
P2 Bounded worker pool, jitter, timeouts (Scheduler)
P3 Content-hash skip on knowledge ingestion (Phase 2)
P4 Hourly watermark-based pattern extraction + feedback(created_at) index (004 / Learning)
P5 entity_status replaces state_snapshots; history in retained metrics (R3-6)
P6 SSE bounded buffers, drop-oldest, heartbeats (Observability)
P7 Docker-only builds; no host go build in deploy (Deploy)
A1 Testing strategy embedded in phase gates + CI coverage gates
A2 Observability section + migration 006
A3, O3, O4 Backup/restore/DR section with drills, RTO/RPO
A4 Init-container migrate/seed, per-file transactions, seed_versions
A5 Config hierarchy (DB-native configuration)
A6 Read-only coexistence + cutover/rollback plan
A7, SA7 Notifier via DB rendezvous
D1 UUIDv7 + slug (R3-5)
D2 attribute_schema JSON Schema validation
D3 entity_type status, deprecate-not-delete, ancestor-aware rules
D4 Atomic counter updates + version optimistic locks
D5, O1 Forward-only migrations + pre-deploy dump + rollback runbook
D6 GET /api/v1/export + byte-stable round-trip test
D7 valid_from/valid_to on relationships
O2 External watchdog cron on apps/105
O5 --no-deps rollout, pinned Postgres container
O6 Per-role health checks wired to compose + watchdog
O7 pmset hardening, OrbStack autostart, update scheduling (R3-13)
M1 Gitea Actions CI, gated webhook
M2 Per-actor token bucket + /executions budget + actuator cooldowns
M3 audit_log covers operator REST mutations with OIDC identity
M4 Per-target circuit breaker
M5 Rotation cadences + expiry signals
M6 Digest-pinned images, govulncheck, distroless
M7 SLO table (Observability)
SA1, SG1 Dual-entity pattern for all cognition objects; hypertable PKs include ts
SA2 Learning cannot write governance; proposes via approvals (Layer map)
SA3 Person/Agent/IdentityProvider entity types (Governance BDD)
SA4 All lifecycles have terminal states + recovery paths
SA5 classifications table (004)
SA6 recommended_action on classification, not signal
SA8 Cluster, ComposeStack, StandaloneServer attrs (BDD)
SA9 timescale image, init-container migrations, DDL/DML user split
SG2, SG3 CAGGs without array_agg; idempotent TimescaleDB DDL
SG4 Graceful shutdown spec (Actuator)
SG5 Per-entity advisory locks
SG6 internal/domain + repository pattern (Repo layout)
SG7 Pattern/skill PATCH endpoints (API)
SG8, SG10 Transactional events + post-commit NOTIFY + in-process bus
SG9 Skill versions as composite PK + change metadata
SG11 Sentinel errors + problem+json mapping
SG13 Context-aware SSH
SG14 Pool sizing: api 15 / scheduler 5 / actuator 5 / learning 3; max_connections=80; alert at 80%
SG15 POST /api/v1/executions resource style
SG16 Cursor pagination + envelope
SG17 sqlc.yaml, module path, CGO off, pgx, distroless, embedded migrations
SG18 Unauthenticated internal-only /healthz + /metrics

Appendix B — initial ADRs to write (Phase 0)

0001 Go + single-binary role packaging · 0002 Postgres+TimescaleDB as the only datastore · 0003 DB-native ontology with YAML seeds · 0004 OpenAPI-first API · 0005 UUIDv7 + slug identity · 0006 learning is proposal-only (no self-authorization) · 0007 threat model + trust zones · 0008 forward-only migrations · 0009 SSE over WebSocket · 0010 Infisical with SOPS DR fallback.