# Plan: Oikos — Docker-based agentic homelab OS on mac-mini **Status:** Done (2026-07-08, Code complete) — Phases 0-6 implemented. All scripts, runbooks, and safeguards in place. Two operational cleanup items remain (apps/105 webhooks + LXC archive, requires operator on Proxmox/Gitea). ## Rev 3 changelog Consolidation: - All rev-2 HIGH/MEDIUM remediations (S1–S10, SA1–SA10, SG1–SG18, A1–A7, O1–O7, D1–D7, P1–P7, M1–M7) are now merged inline. Appendix A maps every finding ID to where it is resolved in this document. New in rev 3 (gaps found in external review): - **R3-1 Ontology inheritance in the meta-schema.** The BDD uses generalization heavily (`ComputeEntity <|-- Machine <|-- ProxmoxHost`), and relationships hang off abstract types (`ComputeEntity provides Service`), but the rev-2 `entity_types` table was flat. Added `parent_type` + `is_abstract`; relationship endpoint validation walks the type hierarchy. - **R3-2 Contract-first API.** `api/openapi.yaml` (OpenAPI 3.1) is the source of truth; Go server stubs generated with `oapi-codegen` (chi router, strict server). Future UIs get a generated TypeScript client. Gin dropped. - **R3-3 API semantics for machines and UIs.** RFC 9457 problem+json errors, uniform list envelope, cursor pagination, `Idempotency-Key` on unsafe POSTs, ETag/If-Match optimistic concurrency, role scopes (operator/viewer/agent), CORS policy, and a `/graph` endpoint for UI visualization. - **R3-4 Single binary, role subcommands.** One `cmd/oikos` binary (`oikos api| scheduler|notifier|all`), one Docker image, roles selected by compose `command` (Loki/Temporal pattern). Guarantees version consistency; trivial local dev. - **R3-5 IDs.** UUIDv7 primary keys (app-generated, time-ordered) + unique human `slug` (`host:hubris`). API accepts either. Resolves D1 properly. - **R3-6 P5 actually fixed.** `state_snapshots` (unbounded growth) replaced by a single-row-per-entity `entity_status` table; health *history* lives in `metric_samples` (already retained/rolled up). - **R3-7 Checks as data.** Probe definitions (`check_defs`) live in the DB and are ontology-attached — adding a probe is an API call, not a code deploy. - **R3-8 Signal dedup + flap suppression + maintenance mode.** Partial unique index guarantees one open signal per (entity, kind); occurrence counting; hold-down on flapping; `maintenance_until` on entities suppresses signals/auto-act. - **R3-9 Executable skill format.** `skills.procedure` is a JSON-schema-validated structure (steps/verify/rollback/params) the actuator can run deterministically; markdown is rendered *from* it for humans. - **R3-10 MCP modernized.** Official `github.com/modelcontextprotocol/go-sdk`, Streamable HTTP transport (SSE-only transport is deprecated in the MCP spec), static bearer token auth. Tools are thin wrappers over the same service layer as REST — one behavior, two protocols. - **R3-11 Ledger simplified.** `ledger_entries` table dropped; `ledger` is a SQL view over `executions ⋈ classifications ⋈ approvals`. One less write path to keep consistent; `audit_log` already covers operator mutations. - **R3-12 Release + rollback made concrete.** Images tagged with git SHA, last 5 kept; rollback = redeploy previous tag + `pg_restore` of pre-deploy dump. - **R3-13 macOS deployment realities + dual-path networking.** Docker-in-VM (OrbStack), no host networking, sleep/auto-restart settings, launch-at-login. Outbound SSH goes direct over the LAN; inbound is mesh-primary with a LAN break-glass binding for the API (rev 3.1, operator decision). Previously silent. - **R3-14 SSE event stream.** Server→client push is all we need; SSE is simpler than WebSocket through Caddy and for browser UIs. WebSocket deferred. - **R3-15 Phases now carry acceptance criteria** ("done when" + verification commands) so an implementing agent knows when to stop. - **R3-16 ADRs.** `docs/adr/` with MADR template; the Decisions table below seeds the initial ADRs. Future architecture changes are recorded, not re-litigated. ## Vision Convert this repo into a **Docker-based agentic homelab OS** written in **Go**. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors: - **Operator** (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix. - **Agent** — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and **learns from every action**. All OS services run in Docker on mac-mini, deployed by git push (Gitea webhook → image build → restart). Designed for mac-mini now with a path to multi-node later (see "Multi-node path"). ## Decisions | Question | Decision | |---|---| | Repo structure | One repo, reorganized (see Repo layout). | | Language | **Go** — compiled, type-safe, small containers, goroutines for concurrent probes. | | Packaging | **Single binary** `oikos` with role subcommands; one Docker image (R3-4). | | API style | **OpenAPI-first** — `api/openapi.yaml` is the contract; oapi-codegen + chi; RFC 9457 errors (R3-2/3). | | Hermes runtime | Docker container, gateway mode. | | Agent → homelab access | Hybrid — restricted SSH key in the **actuator only** now; full actuator gateway in Phase 3. Hermes never holds SSH keys. | | Data storage | PostgreSQL 16 + TimescaleDB. DB is the runtime source of truth. | | Config (inventory/ontology/policy) | **DB-native**; YAML files are seed manifests (bootstrap + DR) with round-trip export. | | IDs | UUIDv7 PK + unique `slug` for humans/API (R3-5). | | Operator interface | Remote Hermes + Matrix now; UIs later on top of the API. | | Deploy | Git push → CI green → Gitea webhook → SHA-tagged image build → compose up. | | `bin/homelab` CLI | Thin Go client generated from the OpenAPI spec. | | Secrets | Migrate SOPS+age → Infisical (one age key kept for DR fallback). | | Notifications | Matrix now, behind a `Notifier` interface; DB is the rendezvous (no service-to-service calls). | | MCP | Same binary/service layer as REST; official Go SDK, Streamable HTTP (R3-10). | | Event stream | SSE (`/api/v1/events/stream`); WebSocket deferred (R3-14). | | Feedback loop | Agent learns from execution; **pattern activation and any autonomy expansion require operator approval** (anti-poisoning). | | Ontology | Developed first; meta-schema supports inheritance + abstract types (R3-1). | | apps/105 | Keeps running (read-only toward shared state) as fallback until cutover. | | ADRs | `docs/adr/` (MADR format); this table seeds ADR-0001…0010 (R3-16). | ## Ontology — the systems model The ontology defines what exists, how things connect, how they change over time, and how the OS learns. It is stored **in the database** (entity types, relationship types, lifecycle definitions). YAML seeds bootstrap it; after that the DB is authoritative and editable via API. ### Design principles 1. **Three layers** — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's behavior + learning). Dependencies flow downward: Cognition depends on Governance depends on Infrastructure. 2. **Everything is an entity** — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, executions, skills) are first-class entities. 3. **Typed relationships with cardinality** — edges carry semantics and are queryable (blast radius, dependency chains, knowledge lookup). 4. **Inheritance is part of the model** — entity types form an is-a hierarchy with abstract types (`ComputeEntity`); relationship endpoint constraints may name abstract types and validation walks the hierarchy (R3-1). 5. **Lifecycles are state machines** — every entity type has a lifecycle with explicit terminal states and (named, code-implemented) transition preconditions. 6. **Policy attaches to the ontology** — risk classes and approval rules link to entity types and actions. 7. **Learning is modeled but never self-authorizing** — executions → feedback → patterns → skills is an explicit, queryable graph, but the learning engine only *proposes* governance changes; the operator approves them (S4/SA2). ### Layer map ```mermaid graph TB cognition -- "observes, acts on, learns about" --> infra governance -- "governs access to" --> infra governance -- "constrains" --> cognition cognition -. "proposes changes (operator approves)" .-> governance subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"] direction LR OBS["Observation\nsignal, check, entity-status"] DEC["Decision\nclassification"] ACT["Action\nexecution, verification"] APPR["Approvals\napproval-request, approval-decision"] KNOW["Knowledge\ndocument, runbook"] LEARN["Learning\nfeedback, pattern, skill"] end subgraph governance["Layer 2 — Governance (who controls what)"] direction LR IDENT["Identity\nperson, agent, identity-provider"] SEC["Secrets\nsecret, key, access-grant"] POL["Policy\nrisk-class, approval-rule, autonomy-setting"] end subgraph infra["Layer 1 — Infrastructure (the managed world)"] direction LR PHYS["Physical\nsite, machine, ups, sensor"] COMP["Compute\nmachine, vm, container\n(lxc, docker)"] NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"] STOR["Storage\nstorage-pool, volume,\nmount, backup-target"] SOFT["Software\nservice, application, cluster,\ncompose-stack, config-repo,\ndeploy-pipeline"] end ``` Note the dashed arrow: the learning engine **cannot write** to governance tables. A validated pattern that would expand autonomy becomes an approval request; only an operator decision changes policy (SA2, S4). ### Block definition diagrams (SysML BDD) Conventions: `«abstract»` = cannot be instantiated; `<|--` generalization; `*--` composition; `o--` aggregation; `-->` association; multiplicities as labeled. **Infrastructure — compute, storage, network:** ```mermaid classDiagram class ComputeEntity { <> +state lifecycle +attributes jsonb } class Machine { +cpu_arch +ram_gb } class VirtualMachine { +vcpus +memory_mb +disk_gb } class Container { <> +runtime } class LXC { +pve_id +rootfs } class DockerContainer { +image +compose_stack } class ProxmoxHost { +pve_version +cluster_member } class StandaloneServer { +hypervisor +provider +control_level } class Workstation { +os +user } class Appliance { +vendor +model } class Hypervisor { +type +version } class Cluster { +quorum +members } class ComposeStack { +path +services } ComputeEntity <|-- Machine ComputeEntity <|-- VirtualMachine ComputeEntity <|-- Container Machine <|-- ProxmoxHost Machine <|-- StandaloneServer Machine <|-- Workstation Machine <|-- Appliance Container <|-- LXC Container <|-- DockerContainer Machine "1" *-- "0..1" Hypervisor : runs Hypervisor "1" o-- "0..*" VirtualMachine : hosts Hypervisor "1" o-- "0..*" Container : hosts ProxmoxHost "0..*" --> "0..1" Cluster : member-of DockerContainer "0..*" --> "0..1" ComposeStack : part-of class StoragePool { +type lvm, zfs, nfs +capacity_gb } class Volume { +name +size_gb } class Mount { +mount_point +options } StoragePool "1" *-- "0..*" Volume : contains ComputeEntity "1" o-- "0..*" Mount : has Mount "0..*" --> "1" Volume : mounts class NetworkInterface { +mac +ip } class Network { <> } class LAN { +subnet } class Mesh { +provider } class VLAN { +tag } ComputeEntity "1" *-- "0..*" NetworkInterface : has NetworkInterface "0..*" --> "1" Network : connects-to Network <|-- LAN Network <|-- Mesh Network <|-- VLAN ``` **Software + services:** ```mermaid classDiagram class Service { +port +health_url +risk_notes } class Application { +version +config } class ConfigRepo { +url +branch } class DeployPipeline { +trigger +target_path } class IngressRoute { +pattern +upstream } class Certificate { +issuer +expires } class DNSZone { +zone } class DNSRecord { +name +record_type +value } ComputeEntity "1" o-- "0..*" Service : provides Service "1" *-- "0..*" Application : runs Service "0..1" --> "0..1" ConfigRepo : configured-by DeployPipeline "0..*" --> "1" Service : deploys-to IngressRoute "0..*" --> "1" Service : routes-to IngressRoute "0..*" --> "0..1" Certificate : secured-by IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by Service "0..*" --> "0..*" Service : depends-on DNSZone "1" *-- "0..*" DNSRecord : contains DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to ``` **Governance — identity (new in rev 2 remediation, kept):** ```mermaid classDiagram class Person { +matrix_id +oidc_sub } class Agent { +provider +model +gateway_port } class IdentityProvider { +issuer +client_id +auth_mode } class Secret { +path +rotation_days } class AccessGrant { +scope +expires } Person "0..1" --> "0..*" Agent : owns IdentityProvider "1" o-- "0..*" Person : authenticates AccessGrant "0..*" --> "1" Secret : grants Agent "0..*" --> "0..*" AccessGrant : holds ``` **Cognition — operations + learning:** ```mermaid classDiagram class CheckDef { +kind +config +interval } class Signal { +kind +severity +state +evidence +occurrences } class Classification { +risk +route +reasoning } class Execution { +status +result +duration_ms +verified } class Feedback { +outcome +observation +lesson } class Pattern { +confidence +evidence_count +status } class Skill { +procedure +version +success_rate } class Approval { +status +ttl } class Document { +title +content +source_path } class Runbook { +steps +risk_class +verification } CheckDef "0..*" --> "1" Entity : checks CheckDef "1" o-- "0..*" Signal : raises Signal "1" --> "0..*" Classification : classified-by Classification "1" --> "0..1" Execution : precedes Execution "1" *-- "0..1" Feedback : produces Feedback "0..*" --> "0..*" Pattern : contributes-to Pattern "0..*" --> "0..1" Skill : informs Skill "0..1" --> "0..*" Classification : guides Execution "0..1" --> "0..1" Approval : requires Agent "0..*" --> "0..*" Execution : performs Person "0..1" --> "0..*" Approval : decides Entity "1" o-- "0..*" Document : documented-by Entity "1" o-- "0..*" Runbook : procedure-for ``` Design notes carried from rev 2 (validated): VMs and LXCs share storage pools via `Mount`→`Volume` on abstract `ComputeEntity`; not every machine is a Proxmox host (fleet today: 2 PVE hosts, 2 workstations, 1 VPS, 2 VMs, 19 LXCs); Docker containers are first-class (the OS models itself); services attach to any compute entity; documents/runbooks attach to any entity via the root abstract type. ### Lifecycles All lifecycles have explicit terminal states and recovery paths (SA4). Transition preconditions are **named checks implemented in Go** and referenced by ID from `lifecycle_defs.transitions` (e.g. `no-inbound-edges`, `backup-verified`) — the DB stores which checks gate a transition; the code implements them. **Infrastructure:** ```mermaid stateDiagram-v2 [*] --> planned : operator creates entity planned --> provisioning : IP reserved, storage chosen, doc stub planned --> destroyed : cancelled provisioning --> active : mesh joined, health answering, doc complete provisioning --> failed : provision failed active --> migrating : preflight + backup verified migrating --> active : post-verify (ingress, mounts checked) migrating --> failed : migration failed failed --> active : recovered failed --> deprecated : written off active --> deprecated : replacement live or role retired deprecated --> active : un-deprecate (replacement failed) deprecated --> destroyed : backups verified, secrets revoked,\ningress removed, zero inbound edges destroyed --> [*] : archaeology entry recorded ``` **Signal:** ```mermaid stateDiagram-v2 [*] --> raised : check fails / agent finding raised --> acknowledged : agent or operator sees it raised --> muted : operator suppresses (TTL) raised --> resolved : condition cleared (auto-resolve) acknowledged --> acting : actuator starts execution acknowledged --> resolved : manual resolve acknowledged --> muted acting --> resolved : action + verification passed acting --> raised : action failed, re-escalated (retry budget left) acting --> failed : permanent failure — needs operator failed --> acknowledged : operator retries muted --> raised : TTL expired and condition persists resolved --> [*] ``` **Execution:** ```mermaid stateDiagram-v2 [*] --> proposed : classification produced an action proposed --> approved : operator approves (if gated) proposed --> auto_approved : risk class allows auto-act proposed --> denied : operator denies approved --> expired : approval TTL ran out approved --> executing auto_approved --> executing executing --> verified : verification passed executing --> failed : execution or verification failed executing --> timed_out executing --> cancelled : operator abort timed_out --> verifying : check if command completed anyway verifying --> verified verifying --> failed failed --> rolled_back : rollback procedure executed failed --> rollback_failed : rollback also failed — page operator verified --> [*] : feedback recorded failed --> [*] : feedback recorded rolled_back --> [*] : feedback recorded rollback_failed --> [*] : feedback recorded cancelled --> [*] denied --> [*] expired --> [*] ``` **Approval:** `pending → approved | denied | expired`; `approved → revoked` (operator changes mind before execution starts). **Pattern:** `hypothesized → validated → active → deprecated`; `hypothesized → invalidated` (disproven, terminal); `active → invalidated` (new evidence contradicts). **`validated → active` requires operator approval** (S4). **Skill:** `drafted → tested → active`; `active → refined → active` (new version); `tested → failed → drafted`; `drafted|active → deprecated`. ### Cognition loop (observe → decide → act → learn) ```mermaid flowchart TB subgraph observe["Observe"] SIG["Signal raised\n(check failed, drift, agent finding)"] end subgraph decide["Decide"] CLASS["Classification\nrisk × blast radius × confidence\n+ recommended action"] SKILL_LOOKUP["Skill lookup\nbest-known procedure for\n(entity type, action)"] CLASS --> SKILL_LOOKUP end subgraph act["Act"] EXEC["Execution\nrun skill procedure\n(or escalate if none)"] VERIFY["Verification"] EXEC --> VERIFY end subgraph learn["Learn"] OUTCOME["Outcome evaluation"] FEEDBACK["Feedback record"] PATTERN["Pattern extraction"] SKILL_REFINE["Skill refinement\n(operator gates activation)"] OUTCOME --> FEEDBACK --> PATTERN --> SKILL_REFINE end SIG --> CLASS SKILL_LOOKUP --> EXEC VERIFY --> OUTCOME SKILL_REFINE -. informs next decision .-> SKILL_LOOKUP CLASS -- needs approval --> APPROVAL["Approval request → Matrix ✅/❌"] APPROVAL -- approved --> EXEC APPROVAL -- denied --> RESOLVE["Resolve signal (denied)"] ``` `recommended_action` lives on the **classification**, not the signal (SA6): checks raise facts; the classifier decides what to do about them. ## Target architecture ### Container stack on mac-mini One image (`oikos:`), three long-running roles + init jobs, three Docker networks as trust boundaries (S10): - `net-front` — Caddy-facing: `api` only. - `net-data` — Postgres + everything that needs it. - `net-ops` — SSH egress: **actuator role only** (Hermes and API have no SSH). ```mermaid graph TB subgraph mac-mini["mac-mini — Docker host (OrbStack), always-on"] subgraph stack["Docker Compose — image oikos:<sha>"] MIGRATE["init: oikos migrate + seed\n(one-shot, DDL user)"] PG["PostgreSQL 16 + TimescaleDB\nentities/relationships • signals\nexecutions • patterns/skills\npolicy • ontology • hypertables"] INF["Infisical (secrets)"] API["oikos api\nREST (OpenAPI) + MCP (streamable HTTP)\npolicy enforcement • audit • events\nSSE stream"] SCHED["oikos scheduler\nchecks (from check_defs) → signals\n+ actuator: classify → execute/escalate\n+ learning engine\nSSH (restricted key) → fleet"] NOTIF["oikos notifier\nMatrix alerts + approval reactions\nDB rendezvous (no RPC)"] HERMES["Hermes agent (gateway :8092)\nMCP client → api\nNO SSH keys"] end DEPLOY["deploy webhook listener\n(HMAC-verified, non-root)"] end subgraph external["External"] CADDY["Caddy (LXC 121)"] GITEA["Gitea (LXC 104) + CI"] MATRIX["Matrix (LXC 118)"] FLEET["hubris / strong / LXCs / VMs"] WS["Workstations (Hermes remote)"] WATCHDOG["Watchdog cron on apps/105\ncurl /healthz → Matrix ping"] end MIGRATE --> PG API --> PG SCHED --> PG NOTIF --> PG HERMES -- "MCP (bearer token)" --> API SCHED -- "SSH (restricted key)" --> FLEET NOTIF --- MATRIX CADDY -- "reverse_proxy + forward-auth" --> API GITEA -- "webhook (HMAC, CI-gated)" --> DEPLOY WS -- "mesh, mTLS/token" --> HERMES WATCHDOG -. "external heartbeat" .-> API ``` Container hardening: read-only root filesystems, non-root users, `restart: always`, `stop_grace_period: 30s`, images pinned by digest, `gcr.io/distroless/static` runtime base, `CGO_ENABLED=0`, `pgx` (pure Go). ### macOS host realities (R3-13) Docker on macOS runs in a lightweight VM (use **OrbStack**: fast, auto-starts at login, stable networking). Consequences the deploy must respect: - **No `network_mode: host`.** All inbound reachability is via published ports. - **Outbound (actuator SSH → hubris/strong/LXCs): direct over the LAN.** Containers reach the LAN via Docker NAT with no special config; SSH provides its own encryption, so there is no reason to route this hop through the mesh. - **Inbound: mesh-primary, LAN break-glass.** NetBird runs on the macOS host. - Primary: Caddy and workstations reach the API and Hermes gateway on ports published on the **mesh IP** (`:8090` API, `:8092` Hermes gateway) — WireGuard-encrypted Caddy→backend hop, mesh membership as a network-level filter, stable addressing, works for roaming workstations. - Break-glass: the **API port is also published on the mac-mini's LAN IP** (`:8090`; give mac-mini a DHCP reservation). Safe because the API authenticates every request itself (OIDC JWT / bearer token — S6); the network you arrive from is defense-in-depth, not the auth. This keeps the control plane reachable from inside the house if NetBird's management plane is down after a reboot, and lets the watchdog test the API independently of the mesh. The Hermes gateway stays **mesh-only** (no LAN binding — it has the weakest application-layer auth story and no break-glass need). - Never bind published ports to `0.0.0.0`; enumerate mesh IP, LAN IP (API only), and localhost explicitly. - Host prep: disable sleep (`sudo pmset -a sleep 0 displaysleep 10`), auto-restart after power failure (`sudo pmset -a autorestart 1`), auto-login enabled so OrbStack starts, macOS auto-updates deferred/scheduled (O7). - Volumes: keep Postgres data on a named Docker volume (VM-native filesystem), not a bind mount — bind mounts cross the VM boundary and are slow. ### Multi-node path (designed for, not built) All roles are stateless; Postgres is the only stateful service. Scaling later means: run `oikos scheduler` on another node pointed at the same DB (checks can be partitioned by a `zone` attribute on `check_defs`); run `oikos api` behind Caddy on N nodes. `SELECT … FOR UPDATE SKIP LOCKED` + advisory locks already make the work queue multi-consumer-safe. Postgres remains the accepted SPOF (mitigated by backup/DR, below), with an upgrade path to streaming replication if ever needed. ## API contract ### Contract-first (R3-2) - `api/openapi.yaml` (OpenAPI 3.1) is the **source of truth**. CI fails if handlers drift from the spec. - Server: `oapi-codegen` strict-server stubs on `chi` + stdlib `net/http`. - Clients: the `homelab` CLI and future web UIs consume generated clients (Go / TypeScript via `openapi-typescript`). The spec is served at `GET /api/v1/openapi.yaml` and human docs at `GET /api/v1/docs` (Redoc/Scalar static page) — a UI developer needs nothing but the running API. ### Conventions (R3-3) - **Versioning:** everything under `/api/v1`. Additive-only within v1 (new fields, new endpoints); breaking changes ship as `/api/v2` side by side with a deprecation window. - **Errors:** RFC 9457 `application/problem+json`: `{"type":"https://oikos.dev/errors/invalid-transition","title":"invalid lifecycle transition","status":409,"detail":"...","instance":"/api/v1/entities/…","errors":[{field,reason}]}`. Domain sentinel errors map centrally: `ErrNotFound→404`, `ErrInvalidTransition→409`, `ErrApprovalRequired→403`, `ErrAutonomyBlocked→403`, `ErrConflict→409`, `ErrCircuitOpen→503`, validation → 422. - **Lists:** uniform envelope `{"items":[…],"next_cursor":"…"}`. Cursor pagination everywhere (`?cursor=&limit=`, default 50, max 200): keyset on `(created_at,id)` for entity-ish tables, on `ts` for hypertables. MCP tools take the same `limit`. - **Idempotency:** unsafe POSTs (`/executions`, `/approvals/{id}/decision`) accept an `Idempotency-Key` header; keys + response snapshots stored 24h; replay returns the original response. Agents retry safely (R3-3). - **Optimistic concurrency:** mutable resources carry a `version`; `GET` returns `ETag`; `PATCH`/`PUT` require `If-Match`, mismatch → 412. UIs can safely edit. - **Timestamps:** `TIMESTAMPTZ` in DB, RFC 3339 UTC on the wire. - **CORS:** config-driven origin allowlist (empty by default; future UI origins added via config, not code). - **Rate limiting (M2):** token bucket per authenticated actor; tighter budget on `/executions`; per-(entity,action) cooldown enforced in the actuator besides. ### AuthN/AuthZ | Caller | Mechanism | Scope | |---|---|---| | Operator (browser/CLI) | Authentik OIDC; Caddy forward-auth **and** JWT validated in API middleware (defense in depth, S6/SA10) | role `operator` (full) or `viewer` (read-only) | | Hermes (MCP) | Static bearer token from Infisical, dedicated Docker network, HMAC on requests (S2) | role `agent` — read tools + `POST /executions` (which is always policy-gated) | | Internal roles (scheduler/notifier) | Direct DB with least-privilege DB users; no API hop | n/a | | `/healthz`, `/metrics` | No auth, no audit; not exposed via Caddy (SG18) | n/a | Roles are claims checked per-route in generated middleware; the spec annotates each operation with its required scope, so future UIs can render capability-aware. ### REST surface (summary; the OpenAPI file is normative) Inventory + ontology: - `GET/POST /api/v1/entities`, `GET/PATCH /api/v1/entities/{id-or-slug}` (PATCH covers attribute edits and lifecycle transitions; transition legality validated against `lifecycle_defs`) - `GET /api/v1/entities/{id}/relations`, `GET /api/v1/graph?root=&depth=&rel_type=` → `{nodes:[…],edges:[…]}` for UI visualization (R3-3) - `GET /api/v1/ontology` (types, relationship types, lifecycles), `POST /api/v1/ontology/entity-types`, `PATCH /api/v1/ontology/entity-types/{name}` (policy-gated `config_mutation`; deprecate-not-delete while instances exist, D3) Operations: - `GET /api/v1/signals`, `POST /api/v1/signals/{id}/ack|resolve|mute` - `GET /api/v1/checks`, `POST /api/v1/checks`, `PATCH /api/v1/checks/{id}` (R3-7) - `GET /api/v1/approvals`, `POST /api/v1/approvals/{id}/decision` - `POST /api/v1/executions` (classify → approval check → enqueue; SG15), `GET /api/v1/executions/{id}`, `POST /api/v1/executions/{id}/cancel` - `GET /api/v1/classifications?signal_id=…` Learning (operator safety valves, SG7): - `GET /api/v1/patterns`, `PATCH /api/v1/patterns/{id}` (activate/invalidate — policy-gated `config_mutation`, audit-logged) - `GET /api/v1/skills`, `GET /api/v1/skills/{id}/versions`, `PATCH /api/v1/skills/{id}` Policy (dual-control, S3): - `GET /api/v1/policy/risk-classes|approval-rules|autonomy` - `PATCH` on any policy resource creates a **meta-approval**; the change applies only after operator approval; before/after hash audit-logged. Knowledge + observability: - `GET /api/v1/knowledge/search?q=`, `GET /api/v1/knowledge/{entity_id}` - `GET /api/v1/metrics?entity_id=&metric=&from=&to=&rollup=raw|1h|1d` (auto-selects resolution by range), `GET /api/v1/trends/{entity_id}` - `GET /api/v1/audit?actor=&entity_id=&action=&correlation_id=&from=&to=` - `GET /api/v1/events?type=&entity_id=&severity=&from=&to=` and `GET /api/v1/events/stream` (SSE; `Last-Event-ID` resume; heartbeat comments; bounded per-subscriber buffers, drop-oldest, P6) - `GET /api/v1/agent-activity`, `GET /api/v1/health` (fleet summary + trends) - `GET /api/v1/export` — regenerates the three seed YAMLs from DB (round-trip tested, D6) ### MCP interface (R3-10) - Official `github.com/modelcontextprotocol/go-sdk`, **Streamable HTTP** transport, mounted at `/mcp` on the same binary, bearer-token auth, dedicated network. - Tools delegate to the identical service layer as REST (one behavior, two protocols): `get_entity`, `list_entities`, `get_relations`, `get_blast_radius`, `search_knowledge`, `get_signal_history`, `get_patterns`, `get_skills`, `request_execution`, `query_metrics`, `get_trend`, `get_audit_trail`, `get_event_timeline`, `get_agent_activity`, `get_health_summary`. - Docs/runbooks additionally exposed as **MCP resources** (URI = entity slug) so Hermes can attach them as context without a tool round-trip. ## DB-native configuration `seeds/ontology.yaml`, `seeds/inventory.yaml`, `seeds/policy.yaml` bootstrap the DB and serve DR; afterwards the DB is authoritative and editable via API. Ingest runs in the one-shot init container (A4): each seed file applies in a single transaction; a `seed_versions` table records applied (file, content-hash) so unchanged seeds are skipped; `GET /api/v1/export` regenerates the YAMLs for commit. Round-trip (seed → DB → export → DB) must be byte-stable — tested in CI. Config hierarchy (A5): compiled defaults → config file → env vars → Infisical (**secrets only** — never plain config). Each role's required keys documented in `docs/operations/config.md`. ## Repo layout ``` / ├── docker-compose.yml ├── Makefile # build, test, lint, generate, deploy targets ├── go.mod / go.sum ├── sqlc.yaml ├── .golangci.yml ├── api/ │ └── openapi.yaml # THE API contract (R3-2) ├── cmd/ │ └── oikos/ # single binary: api | scheduler | notifier | all | │ └── main.go # migrate | seed | export (R3-4) ├── internal/ │ ├── domain/ # pure domain types + state machines + sentinel errors │ │ ├── entity.go signal.go execution.go classification.go │ │ ├── pattern.go skill.go approval.go check.go errors.go │ ├── db/ # pgx pool, sqlc output, repositories (models never escape) │ │ └── queries/ # sqlc SQL │ ├── ontology/ # type hierarchy, validation, graph traversal, seed ingest/export │ ├── httpapi/ # oapi-codegen server impl, middleware (auth, audit, │ │ # idempotency, rate-limit, problem+json mapping), SSE │ ├── mcp/ # MCP server (official SDK) over the same services │ ├── service/ # shared service layer used by httpapi + mcp + loops │ ├── policy/ # classify, approve (tokens), autonomy, meta-approval │ ├── scheduler/ # check runner (check_defs → signals/metrics), dedup, flap │ ├── actuator/ # queue consumer, SSH exec, verify, circuit breaker, locks │ ├── learning/ # feedback, pattern extraction, skill refinement │ ├── notifier/ # interface + matrix impl (DB rendezvous) │ ├── observability/ # slog setup, metrics, audit, events, correlation │ └── config/ ├── migrations/ # golang-migrate, embedded, forward-only (D5/O1) ├── seeds/ # ontology.yaml, inventory.yaml, policy.yaml ├── docs/ │ ├── adr/ # MADR records (R3-16) │ ├── operations/ # backup-restore.md, dr.md, config.md, runbooks │ └── … # narrative docs (ingested into knowledge graph) ├── hermes/ # config.yaml, SOUL.md, skills/homelab-ops/SKILL.md ├── compose/ │ ├── oikos/Dockerfile # one multi-stage Dockerfile for the binary │ └── postgres/ # timescale/timescaledb:2-pg16 config └── scripts/ # migrate-sops.sh, import-legacy.sh, deploy.sh, watchdog.sh ``` Developer experience: `make generate` (oapi-codegen + sqlc), `make test` (unit + testcontainers), `make dev` (`docker compose --profile dev up` with seeded fake data + `oikos all`), `.env.example` committed. ## Database schema Consolidated migrations — all rev-2 fixes applied inline. Forward-only (no `down.sql`; compensating migrations for rollback, plus pre-deploy dumps). Runner: `golang-migrate` via `oikos migrate` in the init container with a DDL-only DB user; runtime roles get DML-only users (SA9). `updated_at` maintained by a shared trigger. ### 001 — Ontology meta-schema (with inheritance, R3-1) ```sql CREATE TABLE lifecycle_defs ( id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', ... states TEXT[] NOT NULL, default_state TEXT NOT NULL, terminal_states TEXT[] NOT NULL DEFAULT '{}', transitions JSONB NOT NULL, -- {"from":{"to":{"requires":["no-inbound-edges",...]}}} created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE entity_types ( name TEXT PRIMARY KEY, -- 'compute-entity', 'machine', 'proxmox-host' parent_type TEXT REFERENCES entity_types(name), -- is-a hierarchy (R3-1) is_abstract BOOLEAN NOT NULL DEFAULT false, -- abstract types can't be instantiated domain TEXT NOT NULL, -- 'physical','compute','network','storage', -- 'software','identity','policy','cognition' layer TEXT NOT NULL CHECK (layer IN ('meta','infrastructure','governance','cognition')), -- 'meta' is reserved for the abstract root type 'entity' description TEXT, lifecycle_id TEXT REFERENCES lifecycle_defs(id), attribute_schema JSONB, -- JSON Schema for entities.attributes (D2) schema_version INTEGER NOT NULL DEFAULT 1, status TEXT NOT NULL DEFAULT 'active', -- 'active'|'deprecated'; no hard delete -- while instances exist (D3) created_at TIMESTAMPTZ NOT NULL DEFAULT now(), updated_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE relationship_types ( name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on' inverse TEXT, source_type TEXT NOT NULL REFERENCES entity_types(name), -- MAY be abstract; target_type TEXT NOT NULL REFERENCES entity_types(name), -- validation walks hierarchy cardinality TEXT NOT NULL CHECK (cardinality IN -- source→target multiplicity ('one-to-one','one-to-many','many-to-one','many-to-many')), description TEXT, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE seed_versions ( -- A4: skip unchanged seed files file TEXT PRIMARY KEY, content_hash TEXT NOT NULL, applied_at TIMESTAMPTZ NOT NULL DEFAULT now() ); ``` Validation semantics (app layer, `internal/ontology`): - Instantiating an `is_abstract` type is rejected. - A relationship `(s, t, type)` is valid iff `type_of(s)` is `source_type` **or a descendant of it** (same for target). The hierarchy is small; resolved in-memory with a cached type tree. - Cardinality enforced by partial unique indexes where expressible (`one-to-many` → `UNIQUE (target_id, type)`; `one-to-one` → unique on both ends) plus app-layer checks for the rest. ### 002 — Entity instances (UUIDv7 + slug, R3-5/D1) ```sql CREATE TABLE entities ( id UUID PRIMARY KEY, -- UUIDv7 generated in Go (time-ordered) slug TEXT NOT NULL UNIQUE, -- 'host:hubris', 'service:caddy' — human/API handle type TEXT NOT NULL REFERENCES entity_types(name), name TEXT NOT NULL, state TEXT, -- lifecycle state attributes JSONB NOT NULL DEFAULT '{}', -- validated against attribute_schema maintenance_until TIMESTAMPTZ, -- R3-8: suppress signals + auto-act while set version INTEGER NOT NULL DEFAULT 1, -- optimistic lock / ETag source created_at TIMESTAMPTZ NOT NULL DEFAULT now(), updated_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (type, name) ); CREATE INDEX idx_entities_type ON entities(type); CREATE INDEX idx_entities_state ON entities(state); CREATE INDEX idx_entities_attrs ON entities USING GIN(attributes); CREATE TABLE relationships ( source_id UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT, target_id UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT, type TEXT NOT NULL REFERENCES relationship_types(name), attributes JSONB, valid_from TIMESTAMPTZ NOT NULL DEFAULT now(), -- D7: temporal edges valid_to TIMESTAMPTZ, -- NULL = current PRIMARY KEY (source_id, target_id, type, valid_from) ); CREATE INDEX idx_rel_source ON relationships(source_id) WHERE valid_to IS NULL; CREATE INDEX idx_rel_target ON relationships(target_id) WHERE valid_to IS NULL; CREATE INDEX idx_rel_type ON relationships(type) WHERE valid_to IS NULL; -- Cycle-safe traversal (P1): path accumulator prevents revisits; depth capped. CREATE OR REPLACE FUNCTION blast_radius(start_id UUID, max_depth INT DEFAULT 3, rel_types TEXT[] DEFAULT NULL) RETURNS TABLE(entity_id UUID, depth INT) AS $$ WITH RECURSIVE walk AS ( SELECT start_id AS entity_id, 0 AS depth, ARRAY[start_id] AS path UNION ALL SELECT r.target_id, w.depth + 1, w.path || r.target_id FROM relationships r JOIN walk w ON r.source_id = w.entity_id WHERE w.depth < LEAST(max_depth, 5) AND r.valid_to IS NULL AND NOT r.target_id = ANY(w.path) AND (rel_types IS NULL OR r.type = ANY(rel_types)) ) SELECT entity_id, MIN(depth) FROM walk GROUP BY entity_id; $$ LANGUAGE sql STABLE; ``` Entities are never hard-deleted while edges exist (`ON DELETE RESTRICT`); decommission is the lifecycle path (`… → destroyed`), and destroyed entities remain as archaeology. **Dual-entity pattern (SA1):** every cognition object (signal, classification, execution, feedback, pattern, skill, approval, check) gets an `entities` row (so the graph is traversable: `triggers`, `produces`, `contributes-to` edges live in `relationships`) **and** a typed table below whose PK references `entities(id)` for indexed querying. ### 003 — Operations (signals, checks, approvals, status) ```sql CREATE TABLE check_defs ( -- R3-7: probes as data entity_id UUID PRIMARY KEY REFERENCES entities(id), target_id UUID REFERENCES entities(id), -- what it checks (NULL + target_type = type-scoped) target_type TEXT REFERENCES entity_types(name), kind TEXT NOT NULL, -- 'http','tcp','disk','cert-expiry','drift','ssh-script' config JSONB NOT NULL DEFAULT '{}', -- validated per-kind JSON Schema interval_s INTEGER NOT NULL DEFAULT 600, timeout_s INTEGER NOT NULL DEFAULT 10, zone TEXT, -- multi-node partitioning later enabled BOOLEAN NOT NULL DEFAULT true, updated_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE signals ( entity_id UUID PRIMARY KEY REFERENCES entities(id), kind TEXT NOT NULL, severity TEXT NOT NULL CHECK (severity IN ('info','warning','critical')), target_entity_id UUID REFERENCES entities(id), check_id UUID REFERENCES check_defs(entity_id), evidence TEXT, likely_cause TEXT, state TEXT NOT NULL DEFAULT 'raised', occurrence_count INTEGER NOT NULL DEFAULT 1, -- R3-8: dedup counting first_seen_at TIMESTAMPTZ NOT NULL DEFAULT now(), last_seen_at TIMESTAMPTZ NOT NULL DEFAULT now(), flap_count INTEGER NOT NULL DEFAULT 0, -- resolve→re-raise cycles hold_down_until TIMESTAMPTZ, -- flap suppression window mute_until TIMESTAMPTZ, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), updated_at TIMESTAMPTZ NOT NULL DEFAULT now() ); -- R3-8: at most ONE open signal per (target, kind) — repeats update the open row CREATE UNIQUE INDEX uq_signals_open ON signals(target_entity_id, kind) WHERE state NOT IN ('resolved','failed'); CREATE INDEX idx_signals_state ON signals(state); CREATE TABLE approvals ( entity_id UUID PRIMARY KEY REFERENCES entities(id), subject_entity_id UUID REFERENCES entities(id), -- entity to act on action TEXT NOT NULL, risk_class TEXT NOT NULL, kind TEXT NOT NULL DEFAULT 'execution', -- 'execution'|'policy-change'|'pattern-activation' payload JSONB, -- e.g. the proposed policy diff status TEXT NOT NULL DEFAULT 'pending', token_hash TEXT, -- S5: single-use HMAC token, stored hashed expires_at TIMESTAMPTZ NOT NULL, decided_at TIMESTAMPTZ, decided_by UUID REFERENCES entities(id), -- person entity created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE entity_status ( -- R3-6: replaces state_snapshots (P5) entity_id UUID PRIMARY KEY REFERENCES entities(id), health TEXT NOT NULL DEFAULT 'unknown', -- healthy|degraded|down|unknown last_check_at TIMESTAMPTZ, details JSONB NOT NULL DEFAULT '{}', updated_at TIMESTAMPTZ NOT NULL DEFAULT now() ); -- health HISTORY is the 'health' metric in metric_samples (retained + rolled up) CREATE TABLE idempotency_keys ( -- R3-3 key TEXT NOT NULL, actor TEXT NOT NULL, request_hash TEXT NOT NULL, response_code INTEGER, response_body JSONB, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (actor, key) ); -- pruned by the scheduler after 24h ``` Approval tokens (S5): token = HMAC(approval_id ‖ subject ‖ action ‖ risk_class ‖ nonce, secret), single-use, stored hashed, TTL-bound; Matrix carries only the approval ID + decision; verification is server-side. ### 004 — Cognition (classifications, executions, learning) ```sql CREATE TABLE classifications ( -- SA5: every autonomous decision persisted entity_id UUID PRIMARY KEY REFERENCES entities(id), signal_entity_id UUID REFERENCES signals(entity_id), target_entity_id UUID REFERENCES entities(id), action TEXT NOT NULL, recommended_action JSONB, -- SA6: lives here, not on the signal risk_class TEXT NOT NULL, route TEXT NOT NULL CHECK (route IN ('auto-act','escalate','hold')), blast_radius UUID[], pattern_confidence REAL, skill_id UUID, -- skill entity matched (if any) autonomy_check TEXT, -- 'allowed' | 'blocked: ' reasoning JSONB NOT NULL, correlation_id TEXT NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE executions ( entity_id UUID PRIMARY KEY REFERENCES entities(id), classification_id UUID REFERENCES classifications(entity_id), signal_entity_id UUID REFERENCES signals(entity_id), target_entity_id UUID REFERENCES entities(id), action TEXT NOT NULL, risk_class TEXT NOT NULL, approval_id UUID REFERENCES approvals(entity_id), agent_id UUID REFERENCES entities(id), skill_id UUID, -- + version pinned at execution time (SG9) skill_version INTEGER, status TEXT NOT NULL DEFAULT 'proposed', result JSONB, duration_ms INTEGER, verified BOOLEAN NOT NULL DEFAULT false, correlation_id TEXT NOT NULL, started_at TIMESTAMPTZ, completed_at TIMESTAMPTZ, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE INDEX idx_exec_target ON executions(target_entity_id); CREATE INDEX idx_exec_status ON executions(status); CREATE TABLE feedback ( entity_id UUID PRIMARY KEY REFERENCES entities(id), execution_id UUID NOT NULL REFERENCES executions(entity_id), outcome TEXT NOT NULL CHECK (outcome IN ('success','failure','partial','unexpected')), observation TEXT, lesson TEXT, unexpected_side_effects TEXT[], tags TEXT[], created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE INDEX idx_feedback_ts ON feedback(created_at); -- P4: watermark scans CREATE TABLE patterns ( entity_id UUID PRIMARY KEY REFERENCES entities(id), applies_type TEXT NOT NULL REFERENCES entity_types(name), action TEXT NOT NULL, pattern TEXT NOT NULL, confidence REAL NOT NULL DEFAULT 0, -- Wilson lower bound, capped by sample size (S4) evidence_count INTEGER NOT NULL DEFAULT 0, success_count INTEGER NOT NULL DEFAULT 0, failure_count INTEGER NOT NULL DEFAULT 0, status TEXT NOT NULL DEFAULT 'hypothesized', quarantined BOOLEAN NOT NULL DEFAULT false, -- S4: anomalous feedback bursts version INTEGER NOT NULL DEFAULT 1, -- D4: optimistic lock last_validated_at TIMESTAMPTZ, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (applies_type, action) ); -- Counters updated atomically: UPDATE … SET evidence_count = evidence_count + 1 (D4) CREATE TABLE skills ( entity_id UUID NOT NULL REFERENCES entities(id), version INTEGER NOT NULL DEFAULT 1, -- SG9: history preserved name TEXT NOT NULL, procedure JSONB NOT NULL, -- R3-9: structured, schema-validated (below) applies_type TEXT REFERENCES entity_types(name), action TEXT NOT NULL, pattern_ids UUID[], status TEXT NOT NULL DEFAULT 'drafted', success_rate REAL, changed_by UUID, -- agent or person entity change_reason TEXT, last_used_at TIMESTAMPTZ, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (entity_id, version) ); ``` **Skill procedure format (R3-9)** — validated by JSON Schema at write time; the actuator executes it deterministically; markdown for humans is rendered from it: ```json { "params_schema": { "type": "object", "properties": { "unit": {"type":"string"} }, "required": ["unit"] }, "steps": [ { "name": "restart unit", "runner": "ssh", "target": "{{ .host }}", "command": "systemctl restart {{ .unit }}", "timeout_s": 60 } ], "verify": [ { "runner": "ssh", "target": "{{ .host }}", "command": "systemctl is-active {{ .unit }}", "expect": { "exit_code": 0, "stdout_contains": "active" }, "retry": { "attempts": 3, "delay_s": 10 } } ], "rollback": [], "expected_duration_s": 30, "known_failure_modes": ["unit masked", "dependency service down"] } ``` Templates are Go `text/template` over validated params; the actuator refuses any command not generated from a stored skill/runbook step (no free-form agent shell). ### 005 — Policy ```sql CREATE TABLE risk_classes ( name TEXT PRIMARY KEY, -- read_only, reversible_low, config_mutation, destructive description TEXT, approval_required TEXT NOT NULL DEFAULT 'none', -- none|operator|operator_confirmed autonomy_allowed BOOLEAN NOT NULL DEFAULT false ); CREATE TABLE approval_rules ( id UUID PRIMARY KEY, entity_type TEXT REFERENCES entity_types(name), -- may be abstract (R3-1) action TEXT NOT NULL, risk_class TEXT NOT NULL REFERENCES risk_classes(name), autonomy_level TEXT NOT NULL DEFAULT 'escalate' CHECK (autonomy_level IN ('auto','escalate','never')), scope_entity UUID REFERENCES entities(id), -- optional per-entity override version INTEGER NOT NULL DEFAULT 1, updated_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (entity_type, action, scope_entity) ); CREATE TABLE autonomy_settings ( key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.' value TEXT NOT NULL, version INTEGER NOT NULL DEFAULT 1, updated_at TIMESTAMPTZ NOT NULL DEFAULT now() ); ``` Rule resolution: most-specific wins (scope_entity > concrete type > ancestor type via the hierarchy). Policy mutations are dual-controlled (S3): the API writes a `policy-change` approval; on operator approval the change applies in one transaction with a before/after hash written to `audit_log`; at startup each role verifies the policy hash against last-known-good and raises a critical signal on mismatch. ### 006 — Observability (TimescaleDB) Hypertable PKs include the time column (SG1); TimescaleDB DDL is idempotent (`if_not_exists => TRUE`, exception-guarded policies — SG3); CAGGs avoid `array_agg` (SG2). ```sql CREATE EXTENSION IF NOT EXISTS timescaledb; CREATE TABLE metric_samples ( ts TIMESTAMPTZ NOT NULL, entity_id UUID NOT NULL, metric TEXT NOT NULL, -- 'health','disk_usage_pct','probe_latency_ms', -- 'api_latency_ms','pattern_confidence',… value DOUBLE PRECISION NOT NULL, tags JSONB NOT NULL DEFAULT '{}' ); SELECT create_hypertable('metric_samples','ts', chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE); CREATE INDEX idx_metrics_entity_ts ON metric_samples(entity_id, ts DESC); CREATE INDEX idx_metrics_metric_ts ON metric_samples(metric, ts DESC); DO $$ BEGIN PERFORM add_retention_policy('metric_samples', INTERVAL '90 days'); EXCEPTION WHEN OTHERS THEN NULL; END $$; CREATE MATERIALIZED VIEW metric_rollups_1h WITH (timescaledb.continuous) AS SELECT time_bucket('1 hour', ts) AS bucket, entity_id, metric, avg(value) AS avg_value, min(value) AS min_value, max(value) AS max_value, count(*) AS sample_count FROM metric_samples GROUP BY bucket, entity_id, metric; -- + metric_rollups_1d identically; refresh policies 1h/1d; rollups kept 1 year CREATE TABLE audit_log ( id BIGINT GENERATED ALWAYS AS IDENTITY, ts TIMESTAMPTZ NOT NULL DEFAULT now(), actor_type TEXT NOT NULL, -- agent|operator|system|scheduler actor_id UUID, -- resolves to a real entity (SA3) action TEXT NOT NULL, entity_id UUID, method TEXT, path TEXT, status_code INTEGER, detail JSONB NOT NULL DEFAULT '{}', source_ip TEXT, correlation_id TEXT, PRIMARY KEY (id, ts) -- SG1 ); SELECT create_hypertable('audit_log','ts', chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE); -- indexes on (actor_type,actor_id,ts), (entity_id,ts), (correlation_id); 365d retention CREATE TABLE events ( id BIGINT GENERATED ALWAYS AS IDENTITY, ts TIMESTAMPTZ NOT NULL DEFAULT now(), type TEXT NOT NULL, -- 'signal.raised','execution.completed',… entity_id UUID, severity TEXT NOT NULL DEFAULT 'info', source TEXT NOT NULL, data JSONB NOT NULL DEFAULT '{}', correlation_id TEXT, PRIMARY KEY (id, ts) ); SELECT create_hypertable('events','ts', chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE); -- 90d retention; NOTIFY trigger fires after commit for the SSE stream (SG10) CREATE TABLE agent_activity ( id BIGINT GENERATED ALWAYS AS IDENTITY, ts TIMESTAMPTZ NOT NULL DEFAULT now(), agent_id UUID NOT NULL, session_id TEXT, activity_type TEXT NOT NULL, -- tool_call|reasoning|decision|mcp_query|escalation tool_name TEXT, entity_id UUID, input_summary TEXT, output_summary TEXT, -- truncated 500 chars duration_ms INTEGER, token_count INTEGER, success BOOLEAN, correlation_id TEXT, PRIMARY KEY (id, ts) ); SELECT create_hypertable('agent_activity','ts', chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE); -- 90d retention -- R3-11: ledger is a VIEW, not a fourth write path CREATE VIEW ledger AS SELECT e.created_at AS ts, e.entity_id AS execution_id, e.target_entity_id, e.action, e.risk_class, e.status, e.verified, c.route, c.reasoning, a.status AS approval_status, a.decided_by, e.agent_id, e.correlation_id FROM executions e LEFT JOIN classifications c ON c.entity_id = e.classification_id LEFT JOIN approvals a ON a.entity_id = e.approval_id; ``` **Retention summary:** metrics 90d raw / 1y rollups; audit 1y; events 90d; agent activity 90d; signals/executions/classifications/feedback permanent (they're the learning corpus and small); `idempotency_keys` 24h (scheduler prune job). ## Control loop ### Scheduler (Observe) - Loads enabled `check_defs`; each check runs on its own interval with jitter; a bounded worker pool (`errgroup.SetLimit`) caps concurrency; per-check timeouts (P2). - Check kinds implemented in Go, config-driven: `http`, `tcp`, `disk` (SSH `df`), `cert-expiry`, `drift` (DB inventory vs live state — mismatches raise `drift` signals with the observed diff as evidence), `ssh-script` (allowlisted). - Every run writes metrics (`health`, `probe_latency_ms`, …) and updates `entity_status` in place (R3-6). - **Signal dedup (R3-8):** failure upserts against the partial unique index — an existing open signal gets `occurrence_count+1`, `last_seen_at=now()`. Recovery auto-resolves. A resolve→re-raise cycle increments `flap_count`; after 3 cycles/1h the signal enters hold-down (no notifications, no auto-act) and a `flapping` meta-signal is raised for the operator. - Entities with `maintenance_until > now()` still get metrics but no signals and are excluded from auto-act. - Housekeeping jobs: daily `pg_dump` + rclone push, idempotency-key prune, feedback watermark advance, secrets-expiry check (M5). ### Actuator (Act) - Consumes signals with `route='auto-act'` classifications via `SELECT … FOR UPDATE SKIP LOCKED`; per-target serialization with `pg_advisory_xact_lock(hashtext(target_entity_id::text))` (SG5). - Executes only stored skill/runbook procedures (R3-9) over SSH with a **restricted key**: dedicated keypair, `command=`/`from=` constrained in `authorized_keys` on targets, mounted read-only into the scheduler/actuator container only (S1). Phase 3 formalizes this as the gateway: all execution flows through `/executions`; Hermes never touches SSH. - Context-aware SSH (SG13): session started in a goroutine, `ctx.Done()` closes session+client to unblock; SSH errors classified (network → retryable/circuit, auth → fatal alert, non-zero exit → failed, timeout → `timed_out → verifying`). - **Circuit breaker per target host** (M4): N consecutive failures → open circuit, exponential backoff, raise `target-unreachable` signal instead of piling failed executions into the learning corpus. - Loop guard: retry budget per (entity, action) from execution history; rate cooldown per (entity, action). - Autonomy kill-switch consulted every pass (`global.auto_act`, `never_auto_act.`). - Graceful shutdown (SG4): `signal.NotifyContext`; stop intake → 30s drain → if an execution is in flight, mark `failed` ("shutdown interrupted") + emit feedback → close pool. Compose sets `stop_grace_period: 30s`. ### Learning engine Concrete algorithm (P4 + S4 guardrails): 1. Hourly, read feedback past the watermark, joined to executions, grouped by `(applies_type, action)`. 2. Update pattern counters atomically; recompute confidence as the **Wilson score lower bound** of success rate (conservative for small N), additionally capped by `min(confidence, evidence_count/5)` so nothing looks confident before 5 samples. 3. Status: `hypothesized` (N<5) → `validated` (N≥5 and confidence ≥ 0.7, emits `pattern.validated` event + notification) → **`active` only via operator PATCH** (policy-gated `config_mutation`). 4. Anomaly quarantine: >10 identical-outcome feedback rows within 1h for one (type, action) → `quarantined=true`, meta-signal for review. 5. Skill refinement: an `active` pattern with confidence > 0.7 drafts/refines a skill (new version row; `changed_by`, `change_reason` recorded). Skills follow their own lifecycle; activation is operator-gated. **No skill ever auto-promotes an action into `destructive` autonomy** — hard-coded, not policy data. 6. Classifier consumption: pattern confidence and skill existence feed the auto-act/escalate score; frequent-failure patterns *lower* confidence. ### Notifier `Notifier` interface (SendAlert / SendApprovalRequest / decision intake); Matrix implementation posts approvals with ✅/❌ reactions and writes decisions **directly to the approvals table** — the DB is the rendezvous, no service-to-service calls, so pending approvals survive restarts of either side (SA7/A7). If the notifier is down, the operator's alternative path is the REST `/approvals` endpoint. ## Security model Threat model (documented in `docs/adr/0007-threat-model.md`): | Boundary | Mechanism | |---|---| | Internet/mesh → API | Caddy (TLS) + Authentik forward-auth **and** in-API OIDC JWT validation (Caddy compromise ≠ API compromise) | | LAN → API (break-glass) | Same in-API auth (OIDC JWT / bearer) — network origin is defense-in-depth, never the auth. Plaintext hop accepted for emergency/watchdog use only; routine traffic uses the mesh | | Workstation → Hermes gateway | Mesh membership (network) + gateway token/mTLS (application); **mesh-only, no LAN binding** | | Hermes → API (MCP) | Dedicated Docker network + static bearer token + HMAC | | Roles → Postgres | Least-privilege DB users (DDL only in init; DML per role), TLS on the Docker network (S7) | | Actuator → fleet | Restricted SSH key (`command=`,`from=`), actuator container only; full gateway in Phase 3 (S1/S10) | | Gitea → deploy | HMAC-signed webhook, localhost/mesh-bound listener, non-root deploy user, CI-gated (S8/M1) | | Learning → policy | Structurally impossible: learning role's DB user has **no write grants** on policy tables; changes route through approvals (S3/S4) | | Approvals | Single-use HMAC tokens, hashed at rest, TTL (S5) | | Secrets | Infisical with machine identities; bootstrap root of trust = master key in mac-mini Keychain, backed up offline; one SOPS age key retained for DR until a restore drill passes (S9) | Rotation cadences (M5): SSH actuator key 6mo, Infisical machine tokens 90d, MCP bearer 90d, webhook HMAC 1y — each with a documented procedure and a scheduler check that raises a signal 2 weeks before expiry. Supply chain (M6): images pinned by digest, `govulncheck` + `golangci-lint` in CI, distroless runtime. ## Observability All capture goes to Postgres (hypertables above) through `internal/observability`: - **Metrics** — check results, API request count/latency, Go runtime stats, learning metrics (pattern confidence, skill success rate, auto-act vs escalation ratio), agent token/tool-call counts. Also exposed as a Prometheus-format `/metrics` endpoint (internal only) so Grafana/Prometheus can attach later without schema work. - **Audit** — middleware on every mutating REST call + MCP tool call + actuator SSH command; policy changes carry before/after hashes; operator identity from OIDC claims (M3). - **Events** — emitted **in the same transaction** as the state change (SG10); post-commit `NOTIFY` feeds the SSE stream; in-process bus covers API-local events (SG8). SSE subscribers get bounded buffers with drop-oldest + heartbeats (P6); delivery is best-effort, history via `GET /events`. - **Correlation** — a `correlation_id` is minted at signal creation (or API request) and propagated via `context.Context` through classification → approval → execution → SSH → verification → feedback, linking audit + events end-to-end. - **Logging** — `slog` JSON to stdout; every line carries `service`, `correlation_id`, `entity_id` where applicable; `debug=true` enables probe payloads/SQL/classification reasoning. - **Agent self-inspection** — MCP tools let Hermes query its own history, trends, audit trail, and efficiency (token usage over time). SLOs (M7): check interval 10min ±1min; API p99 < 200ms; deploy < 5min; alert delivery < 30s; watchdog detection < 5min. ## Operations ### Backup / restore / DR (A3, O3, O4) - **Daily `pg_dump`** (custom format, compressed) + WAL archiving for PITR; pushed off-host to **Proton Drive via rclone** (reuse existing rclone credentials from LXC 132 setup). Retention 30 daily + 12 monthly. Pre-deploy dump before every migration run. - Infisical: native backup + secrets exported to one SOPS-age-encrypted file as fallback (the retained age key is the DR escape hatch). - **Monthly automated restore drill**: scratch container, `pg_restore`, run the export round-trip check, alert on failure. - DR targets: **RTO 4h / RPO 24h**. Cold-start runbook (`docs/operations/dr.md`): fresh machine → OrbStack + Docker → clone repo → restore Infisical → `pg_restore` → `docker compose up -d` → verify `/healthz` + fleet health. ### Watchdog (O2) Cron on apps/105 (outside the stack): every 5min, curl `/healthz` on **both paths** — `http://:8090/healthz` (LAN, tests the API itself) and `http://:8090/healthz` (mesh, tests the path Caddy and workstations use) — plus `pg_isready`; on any failure, post directly to the Matrix webhook naming which path failed (LAN-down = stack problem; mesh-down-LAN-up = NetBird problem). This is the orthogonal "who watches the watcher" channel. ### Deploy, release, rollback (O1, O5, M1, R3-12) - CI (Gitea Actions): `go vet`, `golangci-lint run`, `go test ./... -race -cover` (coverage gates: ≥80% `internal/policy` + `internal/learning`, ≥60% elsewhere), `oapi-codegen`/`sqlc` diff check (generated code committed and clean), `govulncheck`, `docker build`. - Deploy: webhook (HMAC-verified, gated on green CI) → `deploy.sh`: `git pull` → pre-deploy `pg_dump` → build image `oikos:` (multi-stage; **no host `go build`**, P7) → `oikos migrate` init job → `docker compose up -d --no-deps` app roles (Postgres container never recreated on routine deploys) → healthcheck-gated. - Migration compatibility: additive-only per deploy window (new columns nullable); new code tolerates previous schema for one deploy. - Rollback: retag compose to the previous SHA (last 5 images kept) → if the migration was the problem, `pg_restore` the pre-deploy dump → `docker compose up -d`. Runbook in `docs/operations/rollback.md`. - Health checks (O6): API `/healthz` (DB ping); scheduler freshness ("last successful check pass < 15min") exposed via `entity_status` self-row; notifier "last poll < 60s"; Hermes gateway ping — all wired into compose `healthcheck` + watchdog. ### Coexistence + cutover (A6) During coexistence apps/105 is **read-only toward shared state** (its scheduler disabled once the new stack's checks are verified). Cutover: verify traffic on the new stack → disable apps/105 services → retarget Caddy + Gitea webhooks → cleanup. Rollback after cutover = re-enable apps/105 (its JSONL state frozen at cutover; reconciliation = re-import from Postgres if ever needed). ## Phasing — with acceptance criteria (R3-15) Each phase ends with explicit "done when" checks an implementing agent can run. **Phase 0 — Ontology + contract (no service code):** - Finalize entity types (incl. abstract hierarchy, Person/Agent/IdentityProvider, Cluster/ComposeStack), relationship types (with endpoint types + cardinality), lifecycles (all terminal states + named preconditions). - Write `seeds/*.yaml`; write `api/openapi.yaml` v1 for the full REST surface; write ADRs 0001–0010 from the Decisions table. - **Done when:** seeds lint against a meta-schema validator script; `openapi.yaml` passes `redocly lint`; operator has reviewed the lifecycle diagrams + spec. **Phase 1 — Foundation (DB + skeleton):** - Go module `github.com/dtoro/oikos`; `cmd/oikos` skeleton with role subcommands; domain layer + sentinel errors; migrations 001–006; `oikos migrate` + `oikos seed` (idempotent, transactional, `seed_versions`); sqlc + repositories; slog; testcontainers harness; daily backup job + watchdog cron installed. - **Done when:** `make test` green including: seed ingest twice = no-op; export round-trip byte-stable; `blast_radius` correct on a cyclic fixture; abstract-type instantiation rejected; relationship endpoint validation honors inheritance; hypertables + CAGGs + retention created idempotently; legacy `signals/*.jsonl` + `ledger/*.jsonl` imported by `scripts/import-legacy.sh`. **Phase 2 — API:** - oapi-codegen server; auth middleware (OIDC JWT + roles, bearer for agent); problem+json mapping; idempotency; ETag/If-Match; rate limiting; audit middleware; transactional event emitter + SSE stream; knowledge ingestion from `docs/` (content-hash skip, P3); MCP server (official SDK) over the shared service layer; pattern/skill/policy endpoints with dual-control approvals; CI pipeline live. - **Done when:** spec-conformance tests pass (schemathesis or generated-client round trip); `curl /api/v1/entities?type=service` returns the fleet; MCP `list_entities` returns the same data; a mutating call without `If-Match` on a stale version → 412; replayed `Idempotency-Key` returns the cached response; SSE stream shows an `entity.created` event; audit rows carry the OIDC sub. **Phase 3 — Control loop:** - Scheduler with `check_defs` runner, dedup/flap/maintenance logic, metrics + `entity_status`; actuator with restricted SSH key, advisory locks, circuit breaker, retry budgets, graceful shutdown; learning engine per the algorithm above; approval tokens; notifier (Matrix, DB rendezvous). - **Done when:** killing a probed service raises exactly one signal (repeats increment `occurrence_count`); a flapping fixture enters hold-down; a `service-down` signal on a `reversible_low` target auto-restarts, verifies, and the full correlation chain (signal → classification → execution → SSH audit → feedback) is queryable by one `correlation_id`; a `destructive` action produces a Matrix approval whose ✅ token is single-use; after 5 successful executions a pattern reaches `validated` and **stays there** until operator PATCH; kill-switch `global.auto_act=off` forces escalation. **Phase 4 — Agent (Hermes):** - Hermes container (gateway :8092, mesh-published, token/mTLS), homelab skills, MCP wiring, agent-activity logging. No SSH keys in this container. - **Done when:** from a workstation over mesh, Hermes answers "what depends on authentik?" via `get_blast_radius`, requests an execution that routes through `/executions` policy gating, and its tool calls appear in `agent_activity`. **Phase 5 — Secrets (Infisical):** - Infisical up; SOPS migrated; services on machine identities; rotation checks; SOPS-age DR fallback exported. - **Done when:** no service reads SOPS at runtime; restore drill of Infisical backup passes; rotation runbooks written. **Phase 6 — Deploy + cutover:** - Full pipeline (CI-gated webhook, SHA images, init migrate, healthcheck rollout); Caddy re-point (`mcp.hubris.network`, `oikos.hubris.network` → mac-mini mesh :8090); end-to-end verification below; apps/105 disabled per cutover checklist; first monthly restore drill executed. - **Done when:** all 14 verification checks pass; watchdog alert fires when the API is stopped manually; rollback drill (previous SHA + pg_restore) rehearsed once. ## Verification (end to end) 1. **Ontology:** `SELECT * FROM entity_types` shows the 3-layer hierarchy incl. abstract types; lifecycle defs match the diagrams; abstract instantiation is rejected via API (422). 2. **DB:** migrations 001–006 apply idempotently; seeds ingest; re-ingest is a no-op. 3. **API:** REST + MCP return identical data for the fleet; problem+json on errors; pagination envelope everywhere. 4. **Scheduler:** check pass writes metrics + `entity_status`; one open signal per (entity, kind) under repeated failure. 5. **Actuator:** classify → auto-act or escalate → execute → verify → feedback, all correlation-linked in audit + events. 6. **Learning:** after N=5 similar executions a pattern is `validated` with a Wilson-bounded confidence; operator PATCH activates it; a skill version appears. 7. **Classifier + autonomy:** high-confidence pattern → auto-act on `reversible_low`; kill-switch off → always escalate; `never_auto_act.` honored. 8. **Hermes:** remote workstation session; MCP tools work; activity logged. 9. **Secrets:** services fetch from Infisical; SOPS retired (except DR key). 10. **Deploy:** push → CI → webhook → SHA image → migrate init → rolling restart; `deploy.triggered`/`deploy.completed` events emitted. 11. **Knowledge:** `search_knowledge("caddy")` returns docs edged to `service:caddy`. 12. **Observability:** metrics/trends/audit/health endpoints return correct shapes; Grafana can read `metric_samples` directly. 13. **Correlation tracing:** one `correlation_id` reconstructs signal → classification → execution → SSH → verification → feedback. 14. **Cutover:** apps/105 stopped; watchdog still alive; production traffic served solely by the Docker stack. ## Risks / trade-offs - **Go rewrite** (~4,400 Python lines replaced) — logic carries over 1:1 (see reuse table); mitigated by the phase gates and the contract-first spec. - **Postgres SPOF** — accepted; mitigated by backups/PITR/DR drills; replication is the future path. - **Learning cold start** — by design: the agent escalates everything until patterns validate and the operator activates them; trust is earned. - **Single-host mac-mini** — watchdog + pmset hardening + documented cold start; multi-node path exists when it matters. - **Infisical bootstrap** — SOPS fallback retained until a restore drill passes. ### Reuse map (Python → Go) | Existing | Becomes | |---|---| | `oikos/decide.py` | `internal/policy/classify.go` (+ pattern/skill inputs) | | `oikos/signal.py` | `internal/scheduler` signal lifecycle (DB-backed, dedup added) | | `oikos/approve.py` | `internal/policy/approve.go` (tokens) + `internal/notifier/matrix.go` | | `oikos/ledger.py` | `executions` + `ledger` view + `audit_log` | | `oikos/policy.py` + `policy.yaml` | `internal/policy` + `seeds/policy.yaml` → tables | | `oikos/drift.py` | `drift` check kind | | `oikos/relations.py` | `internal/ontology/graph.go` (SQL traversal) | | `oikos/report.py` | report endpoints over DB | | `mcp/server.py` | `internal/mcp` (official SDK) | | `bin/homelab` | generated OpenAPI Go client | | `oikos/scheduler.py` | `internal/scheduler` (goroutine worker pool, checks-as-data) | | *(new)* | `internal/learning`, `internal/observability`, `internal/httpapi` | ## Out of scope (for now) - Web UI (the OpenAPI contract + SSE + `/graph` endpoint are built for it; the UI itself comes later). - Multi-node deployment (designed for; see "Multi-node path"). - Vector embeddings / semantic search (Postgres FTS now; `pgvector` is a schema-only addition later). - WebSocket stream (SSE covers current needs). - LLM-assisted skill extraction (patterns are statistical for now). --- ## Appendix A — audit resolution ledger Every rev-2 finding and where rev 3 resolves it. (Full finding text in git history, rev 2 of this file.) | Finding | Resolution | |---|---| | S1, S10 | Security model — restricted SSH key in actuator only; network trust zones; Hermes keyless | | S2 | MCP bearer token + dedicated network (API contract / Security) | | S3 | Policy dual-control + hash audit + startup self-check (005 / Security) | | S4 | Operator-gated pattern activation, Wilson + N/5 cap, quarantine, no destructive auto-promotion (Learning engine) | | S5 | Single-use HMAC approval tokens, hashed (003) | | S6, SA10 | In-API OIDC JWT validation; Caddy = explicit trust root (AuthN/AuthZ) | | S7 | TLS to Postgres; per-role DB users (Security) | | S8 | HMAC webhook, non-root deploy, CI gate (Deploy) | | S9 | Infisical bootstrap root of trust + SOPS DR fallback (Security / Phase 5) | | P1 | Cycle-safe `blast_radius` with path accumulator + depth cap (002) | | P2 | Bounded worker pool, jitter, timeouts (Scheduler) | | P3 | Content-hash skip on knowledge ingestion (Phase 2) | | P4 | Hourly watermark-based pattern extraction + `feedback(created_at)` index (004 / Learning) | | P5 | `entity_status` replaces `state_snapshots`; history in retained metrics (R3-6) | | P6 | SSE bounded buffers, drop-oldest, heartbeats (Observability) | | P7 | Docker-only builds; no host `go build` in deploy (Deploy) | | A1 | Testing strategy embedded in phase gates + CI coverage gates | | A2 | Observability section + migration 006 | | A3, O3, O4 | Backup/restore/DR section with drills, RTO/RPO | | A4 | Init-container migrate/seed, per-file transactions, `seed_versions` | | A5 | Config hierarchy (DB-native configuration) | | A6 | Read-only coexistence + cutover/rollback plan | | A7, SA7 | Notifier via DB rendezvous | | D1 | UUIDv7 + slug (R3-5) | | D2 | `attribute_schema` JSON Schema validation | | D3 | entity_type `status`, deprecate-not-delete, ancestor-aware rules | | D4 | Atomic counter updates + `version` optimistic locks | | D5, O1 | Forward-only migrations + pre-deploy dump + rollback runbook | | D6 | `GET /api/v1/export` + byte-stable round-trip test | | D7 | `valid_from`/`valid_to` on relationships | | O2 | External watchdog cron on apps/105 | | O5 | `--no-deps` rollout, pinned Postgres container | | O6 | Per-role health checks wired to compose + watchdog | | O7 | pmset hardening, OrbStack autostart, update scheduling (R3-13) | | M1 | Gitea Actions CI, gated webhook | | M2 | Per-actor token bucket + `/executions` budget + actuator cooldowns | | M3 | `audit_log` covers operator REST mutations with OIDC identity | | M4 | Per-target circuit breaker | | M5 | Rotation cadences + expiry signals | | M6 | Digest-pinned images, `govulncheck`, distroless | | M7 | SLO table (Observability) | | SA1, SG1 | Dual-entity pattern for all cognition objects; hypertable PKs include `ts` | | SA2 | Learning cannot write governance; proposes via approvals (Layer map) | | SA3 | Person/Agent/IdentityProvider entity types (Governance BDD) | | SA4 | All lifecycles have terminal states + recovery paths | | SA5 | `classifications` table (004) | | SA6 | `recommended_action` on classification, not signal | | SA8 | Cluster, ComposeStack, StandaloneServer attrs (BDD) | | SA9 | timescale image, init-container migrations, DDL/DML user split | | SG2, SG3 | CAGGs without `array_agg`; idempotent TimescaleDB DDL | | SG4 | Graceful shutdown spec (Actuator) | | SG5 | Per-entity advisory locks | | SG6 | `internal/domain` + repository pattern (Repo layout) | | SG7 | Pattern/skill PATCH endpoints (API) | | SG8, SG10 | Transactional events + post-commit NOTIFY + in-process bus | | SG9 | Skill versions as composite PK + change metadata | | SG11 | Sentinel errors + problem+json mapping | | SG13 | Context-aware SSH | | SG14 | Pool sizing: api 15 / scheduler 5 / actuator 5 / learning 3; `max_connections=80`; alert at 80% | | SG15 | `POST /api/v1/executions` resource style | | SG16 | Cursor pagination + envelope | | SG17 | sqlc.yaml, module path, CGO off, pgx, distroless, embedded migrations | | SG18 | Unauthenticated internal-only `/healthz` + `/metrics` | ## Appendix B — initial ADRs to write (Phase 0) 0001 Go + single-binary role packaging · 0002 Postgres+TimescaleDB as the only datastore · 0003 DB-native ontology with YAML seeds · 0004 OpenAPI-first API · 0005 UUIDv7 + slug identity · 0006 learning is proposal-only (no self-authorization) · 0007 threat model + trust zones · 0008 forward-only migrations · 0009 SSE over WebSocket · 0010 Infisical with SOPS DR fallback.