- seeds/ontology.yaml: 59 entity types (5 abstract, is-a hierarchy), 46 relationship types with cardinality, 6 lifecycles with terminal states and named precondition checks - seeds/inventory.yaml: 110 entities / 142 relationships translated from legacy inventory.yaml (fleet, services, ingress, storage, governance, archaeology); thin spots marked for backfill - seeds/policy.yaml: 4 risk classes, 27 approval rules (hierarchy-aware, per-entity overrides), autonomy kill-switch off (cold start) - api/openapi.yaml: full v1 REST contract (40 paths), RFC 9457 errors, cursor pagination, idempotency, ETag/If-Match, scopes; redocly-clean - docs/adr/0001-0010: initial architecture decision records - scripts/validate-seeds.py: Phase 0 gate — hierarchy, lifecycles, endpoints, cardinality, policy cross-refs (0 errors) - plan: layer CHECK gains 'meta' (root type), cardinality gains 'many-to-one' Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
77 KiB
Plan: Oikos — Docker-based agentic homelab OS on mac-mini
Status: Planned (2026-07-06, rev 3) — supersedes rev 2. Rev 3 consolidates the rev-2 audit + remediation layers into one self-consistent spec (no more "read the migration, then read the fix section") and closes newly found gaps. This document is the single source of truth for implementation; an agent should be able to implement phase by phase from this file alone.
Rev 3 changelog
Consolidation:
- All rev-2 HIGH/MEDIUM remediations (S1–S10, SA1–SA10, SG1–SG18, A1–A7, O1–O7, D1–D7, P1–P7, M1–M7) are now merged inline. Appendix A maps every finding ID to where it is resolved in this document.
New in rev 3 (gaps found in external review):
- R3-1 Ontology inheritance in the meta-schema. The BDD uses generalization
heavily (
ComputeEntity <|-- Machine <|-- ProxmoxHost), and relationships hang off abstract types (ComputeEntity provides Service), but the rev-2entity_typestable was flat. Addedparent_type+is_abstract; relationship endpoint validation walks the type hierarchy. - R3-2 Contract-first API.
api/openapi.yaml(OpenAPI 3.1) is the source of truth; Go server stubs generated withoapi-codegen(chi router, strict server). Future UIs get a generated TypeScript client. Gin dropped. - R3-3 API semantics for machines and UIs. RFC 9457 problem+json errors, uniform
list envelope, cursor pagination,
Idempotency-Keyon unsafe POSTs, ETag/If-Match optimistic concurrency, role scopes (operator/viewer/agent), CORS policy, and a/graphendpoint for UI visualization. - R3-4 Single binary, role subcommands. One
cmd/oikosbinary (oikos api| scheduler|notifier|all), one Docker image, roles selected by composecommand(Loki/Temporal pattern). Guarantees version consistency; trivial local dev. - R3-5 IDs. UUIDv7 primary keys (app-generated, time-ordered) + unique human
slug(host:hubris). API accepts either. Resolves D1 properly. - R3-6 P5 actually fixed.
state_snapshots(unbounded growth) replaced by a single-row-per-entityentity_statustable; health history lives inmetric_samples(already retained/rolled up). - R3-7 Checks as data. Probe definitions (
check_defs) live in the DB and are ontology-attached — adding a probe is an API call, not a code deploy. - R3-8 Signal dedup + flap suppression + maintenance mode. Partial unique index
guarantees one open signal per (entity, kind); occurrence counting; hold-down on
flapping;
maintenance_untilon entities suppresses signals/auto-act. - R3-9 Executable skill format.
skills.procedureis a JSON-schema-validated structure (steps/verify/rollback/params) the actuator can run deterministically; markdown is rendered from it for humans. - R3-10 MCP modernized. Official
github.com/modelcontextprotocol/go-sdk, Streamable HTTP transport (SSE-only transport is deprecated in the MCP spec), static bearer token auth. Tools are thin wrappers over the same service layer as REST — one behavior, two protocols. - R3-11 Ledger simplified.
ledger_entriestable dropped;ledgeris a SQL view overexecutions ⋈ classifications ⋈ approvals. One less write path to keep consistent;audit_logalready covers operator mutations. - R3-12 Release + rollback made concrete. Images tagged with git SHA, last 5
kept; rollback = redeploy previous tag +
pg_restoreof pre-deploy dump. - R3-13 macOS deployment realities + dual-path networking. Docker-in-VM (OrbStack), no host networking, sleep/auto-restart settings, launch-at-login. Outbound SSH goes direct over the LAN; inbound is mesh-primary with a LAN break-glass binding for the API (rev 3.1, operator decision). Previously silent.
- R3-14 SSE event stream. Server→client push is all we need; SSE is simpler than WebSocket through Caddy and for browser UIs. WebSocket deferred.
- R3-15 Phases now carry acceptance criteria ("done when" + verification commands) so an implementing agent knows when to stop.
- R3-16 ADRs.
docs/adr/with MADR template; the Decisions table below seeds the initial ADRs. Future architecture changes are recorded, not re-litigated.
Vision
Convert this repo into a Docker-based agentic homelab OS written in Go. The OS is a set of containerized services that manage the homelab autonomously, with the operator in control. Two actors:
- Operator (dtoro) — owns the homelab, expresses intent ("install X", "restart Y"), approves destructive actions. Connects from any workstation via remote Hermes or Matrix.
- Agent — Hermes core + custom homelab skills, running in Docker. Executes orders, monitors the lab, escalates when unsure, and learns from every action.
All OS services run in Docker on mac-mini, deployed by git push (Gitea webhook → image build → restart). Designed for mac-mini now with a path to multi-node later (see "Multi-node path").
Decisions
| Question | Decision |
|---|---|
| Repo structure | One repo, reorganized (see Repo layout). |
| Language | Go — compiled, type-safe, small containers, goroutines for concurrent probes. |
| Packaging | Single binary oikos with role subcommands; one Docker image (R3-4). |
| API style | OpenAPI-first — api/openapi.yaml is the contract; oapi-codegen + chi; RFC 9457 errors (R3-2/3). |
| Hermes runtime | Docker container, gateway mode. |
| Agent → homelab access | Hybrid — restricted SSH key in the actuator only now; full actuator gateway in Phase 3. Hermes never holds SSH keys. |
| Data storage | PostgreSQL 16 + TimescaleDB. DB is the runtime source of truth. |
| Config (inventory/ontology/policy) | DB-native; YAML files are seed manifests (bootstrap + DR) with round-trip export. |
| IDs | UUIDv7 PK + unique slug for humans/API (R3-5). |
| Operator interface | Remote Hermes + Matrix now; UIs later on top of the API. |
| Deploy | Git push → CI green → Gitea webhook → SHA-tagged image build → compose up. |
bin/homelab CLI |
Thin Go client generated from the OpenAPI spec. |
| Secrets | Migrate SOPS+age → Infisical (one age key kept for DR fallback). |
| Notifications | Matrix now, behind a Notifier interface; DB is the rendezvous (no service-to-service calls). |
| MCP | Same binary/service layer as REST; official Go SDK, Streamable HTTP (R3-10). |
| Event stream | SSE (/api/v1/events/stream); WebSocket deferred (R3-14). |
| Feedback loop | Agent learns from execution; pattern activation and any autonomy expansion require operator approval (anti-poisoning). |
| Ontology | Developed first; meta-schema supports inheritance + abstract types (R3-1). |
| apps/105 | Keeps running (read-only toward shared state) as fallback until cutover. |
| ADRs | docs/adr/ (MADR format); this table seeds ADR-0001…0010 (R3-16). |
Ontology — the systems model
The ontology defines what exists, how things connect, how they change over time, and how the OS learns. It is stored in the database (entity types, relationship types, lifecycle definitions). YAML seeds bootstrap it; after that the DB is authoritative and editable via API.
Design principles
- Three layers — Infrastructure (the managed world), Governance (who controls what), Cognition (the OS's behavior + learning). Dependencies flow downward: Cognition depends on Governance depends on Infrastructure.
- Everything is an entity — if it can break, be changed, or hold data, it has an entity type and edges. The OS's own objects (signals, executions, skills) are first-class entities.
- Typed relationships with cardinality — edges carry semantics and are queryable (blast radius, dependency chains, knowledge lookup).
- Inheritance is part of the model — entity types form an is-a hierarchy with
abstract types (
ComputeEntity); relationship endpoint constraints may name abstract types and validation walks the hierarchy (R3-1). - Lifecycles are state machines — every entity type has a lifecycle with explicit terminal states and (named, code-implemented) transition preconditions.
- Policy attaches to the ontology — risk classes and approval rules link to entity types and actions.
- Learning is modeled but never self-authorizing — executions → feedback → patterns → skills is an explicit, queryable graph, but the learning engine only proposes governance changes; the operator approves them (S4/SA2).
Layer map
graph TB
cognition -- "observes, acts on, learns about" --> infra
governance -- "governs access to" --> infra
governance -- "constrains" --> cognition
cognition -. "proposes changes (operator approves)" .-> governance
subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"]
direction LR
OBS["Observation\nsignal, check, entity-status"]
DEC["Decision\nclassification"]
ACT["Action\nexecution, verification"]
APPR["Approvals\napproval-request, approval-decision"]
KNOW["Knowledge\ndocument, runbook"]
LEARN["Learning\nfeedback, pattern, skill"]
end
subgraph governance["Layer 2 — Governance (who controls what)"]
direction LR
IDENT["Identity\nperson, agent, identity-provider"]
SEC["Secrets\nsecret, key, access-grant"]
POL["Policy\nrisk-class, approval-rule, autonomy-setting"]
end
subgraph infra["Layer 1 — Infrastructure (the managed world)"]
direction LR
PHYS["Physical\nsite, machine, ups, sensor"]
COMP["Compute\nmachine, vm, container\n(lxc, docker)"]
NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"]
STOR["Storage\nstorage-pool, volume,\nmount, backup-target"]
SOFT["Software\nservice, application, cluster,\ncompose-stack, config-repo,\ndeploy-pipeline"]
end
Note the dashed arrow: the learning engine cannot write to governance tables. A validated pattern that would expand autonomy becomes an approval request; only an operator decision changes policy (SA2, S4).
Block definition diagrams (SysML BDD)
Conventions: «abstract» = cannot be instantiated; <|-- generalization; *--
composition; o-- aggregation; --> association; multiplicities as labeled.
Infrastructure — compute, storage, network:
classDiagram
class ComputeEntity {
<<abstract>>
+state lifecycle
+attributes jsonb
}
class Machine { +cpu_arch +ram_gb }
class VirtualMachine { +vcpus +memory_mb +disk_gb }
class Container { <<abstract>> +runtime }
class LXC { +pve_id +rootfs }
class DockerContainer { +image +compose_stack }
class ProxmoxHost { +pve_version +cluster_member }
class StandaloneServer { +hypervisor +provider +control_level }
class Workstation { +os +user }
class Appliance { +vendor +model }
class Hypervisor { +type +version }
class Cluster { +quorum +members }
class ComposeStack { +path +services }
ComputeEntity <|-- Machine
ComputeEntity <|-- VirtualMachine
ComputeEntity <|-- Container
Machine <|-- ProxmoxHost
Machine <|-- StandaloneServer
Machine <|-- Workstation
Machine <|-- Appliance
Container <|-- LXC
Container <|-- DockerContainer
Machine "1" *-- "0..1" Hypervisor : runs
Hypervisor "1" o-- "0..*" VirtualMachine : hosts
Hypervisor "1" o-- "0..*" Container : hosts
ProxmoxHost "0..*" --> "0..1" Cluster : member-of
DockerContainer "0..*" --> "0..1" ComposeStack : part-of
class StoragePool { +type lvm, zfs, nfs +capacity_gb }
class Volume { +name +size_gb }
class Mount { +mount_point +options }
StoragePool "1" *-- "0..*" Volume : contains
ComputeEntity "1" o-- "0..*" Mount : has
Mount "0..*" --> "1" Volume : mounts
class NetworkInterface { +mac +ip }
class Network { <<abstract>> }
class LAN { +subnet }
class Mesh { +provider }
class VLAN { +tag }
ComputeEntity "1" *-- "0..*" NetworkInterface : has
NetworkInterface "0..*" --> "1" Network : connects-to
Network <|-- LAN
Network <|-- Mesh
Network <|-- VLAN
Software + services:
classDiagram
class Service { +port +health_url +risk_notes }
class Application { +version +config }
class ConfigRepo { +url +branch }
class DeployPipeline { +trigger +target_path }
class IngressRoute { +pattern +upstream }
class Certificate { +issuer +expires }
class DNSZone { +zone }
class DNSRecord { +name +record_type +value }
ComputeEntity "1" o-- "0..*" Service : provides
Service "1" *-- "0..*" Application : runs
Service "0..1" --> "0..1" ConfigRepo : configured-by
DeployPipeline "0..*" --> "1" Service : deploys-to
IngressRoute "0..*" --> "1" Service : routes-to
IngressRoute "0..*" --> "0..1" Certificate : secured-by
IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by
Service "0..*" --> "0..*" Service : depends-on
DNSZone "1" *-- "0..*" DNSRecord : contains
DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to
Governance — identity (new in rev 2 remediation, kept):
classDiagram
class Person { +matrix_id +oidc_sub }
class Agent { +provider +model +gateway_port }
class IdentityProvider { +issuer +client_id +auth_mode }
class Secret { +path +rotation_days }
class AccessGrant { +scope +expires }
Person "0..1" --> "0..*" Agent : owns
IdentityProvider "1" o-- "0..*" Person : authenticates
AccessGrant "0..*" --> "1" Secret : grants
Agent "0..*" --> "0..*" AccessGrant : holds
Cognition — operations + learning:
classDiagram
class CheckDef { +kind +config +interval }
class Signal { +kind +severity +state +evidence +occurrences }
class Classification { +risk +route +reasoning }
class Execution { +status +result +duration_ms +verified }
class Feedback { +outcome +observation +lesson }
class Pattern { +confidence +evidence_count +status }
class Skill { +procedure +version +success_rate }
class Approval { +status +ttl }
class Document { +title +content +source_path }
class Runbook { +steps +risk_class +verification }
CheckDef "0..*" --> "1" Entity : checks
CheckDef "1" o-- "0..*" Signal : raises
Signal "1" --> "0..*" Classification : classified-by
Classification "1" --> "0..1" Execution : precedes
Execution "1" *-- "0..1" Feedback : produces
Feedback "0..*" --> "0..*" Pattern : contributes-to
Pattern "0..*" --> "0..1" Skill : informs
Skill "0..1" --> "0..*" Classification : guides
Execution "0..1" --> "0..1" Approval : requires
Agent "0..*" --> "0..*" Execution : performs
Person "0..1" --> "0..*" Approval : decides
Entity "1" o-- "0..*" Document : documented-by
Entity "1" o-- "0..*" Runbook : procedure-for
Design notes carried from rev 2 (validated): VMs and LXCs share storage pools via
Mount→Volume on abstract ComputeEntity; not every machine is a Proxmox host
(fleet today: 2 PVE hosts, 2 workstations, 1 VPS, 2 VMs, 19 LXCs); Docker containers
are first-class (the OS models itself); services attach to any compute entity;
documents/runbooks attach to any entity via the root abstract type.
Lifecycles
All lifecycles have explicit terminal states and recovery paths (SA4). Transition
preconditions are named checks implemented in Go and referenced by ID from
lifecycle_defs.transitions (e.g. no-inbound-edges, backup-verified) — the DB
stores which checks gate a transition; the code implements them.
Infrastructure:
stateDiagram-v2
[*] --> planned : operator creates entity
planned --> provisioning : IP reserved, storage chosen, doc stub
planned --> destroyed : cancelled
provisioning --> active : mesh joined, health answering, doc complete
provisioning --> failed : provision failed
active --> migrating : preflight + backup verified
migrating --> active : post-verify (ingress, mounts checked)
migrating --> failed : migration failed
failed --> active : recovered
failed --> deprecated : written off
active --> deprecated : replacement live or role retired
deprecated --> active : un-deprecate (replacement failed)
deprecated --> destroyed : backups verified, secrets revoked,\ningress removed, zero inbound edges
destroyed --> [*] : archaeology entry recorded
Signal:
stateDiagram-v2
[*] --> raised : check fails / agent finding
raised --> acknowledged : agent or operator sees it
raised --> muted : operator suppresses (TTL)
raised --> resolved : condition cleared (auto-resolve)
acknowledged --> acting : actuator starts execution
acknowledged --> resolved : manual resolve
acknowledged --> muted
acting --> resolved : action + verification passed
acting --> raised : action failed, re-escalated (retry budget left)
acting --> failed : permanent failure — needs operator
failed --> acknowledged : operator retries
muted --> raised : TTL expired and condition persists
resolved --> [*]
Execution:
stateDiagram-v2
[*] --> proposed : classification produced an action
proposed --> approved : operator approves (if gated)
proposed --> auto_approved : risk class allows auto-act
proposed --> denied : operator denies
approved --> expired : approval TTL ran out
approved --> executing
auto_approved --> executing
executing --> verified : verification passed
executing --> failed : execution or verification failed
executing --> timed_out
executing --> cancelled : operator abort
timed_out --> verifying : check if command completed anyway
verifying --> verified
verifying --> failed
failed --> rolled_back : rollback procedure executed
failed --> rollback_failed : rollback also failed — page operator
verified --> [*] : feedback recorded
failed --> [*] : feedback recorded
rolled_back --> [*] : feedback recorded
rollback_failed --> [*] : feedback recorded
cancelled --> [*]
denied --> [*]
expired --> [*]
Approval: pending → approved | denied | expired; approved → revoked
(operator changes mind before execution starts).
Pattern: hypothesized → validated → active → deprecated; hypothesized → invalidated (disproven, terminal); active → invalidated (new evidence
contradicts). validated → active requires operator approval (S4).
Skill: drafted → tested → active; active → refined → active (new version);
tested → failed → drafted; drafted|active → deprecated.
Cognition loop (observe → decide → act → learn)
flowchart TB
subgraph observe["Observe"]
SIG["Signal raised\n(check failed, drift, agent finding)"]
end
subgraph decide["Decide"]
CLASS["Classification\nrisk × blast radius × confidence\n+ recommended action"]
SKILL_LOOKUP["Skill lookup\nbest-known procedure for\n(entity type, action)"]
CLASS --> SKILL_LOOKUP
end
subgraph act["Act"]
EXEC["Execution\nrun skill procedure\n(or escalate if none)"]
VERIFY["Verification"]
EXEC --> VERIFY
end
subgraph learn["Learn"]
OUTCOME["Outcome evaluation"]
FEEDBACK["Feedback record"]
PATTERN["Pattern extraction"]
SKILL_REFINE["Skill refinement\n(operator gates activation)"]
OUTCOME --> FEEDBACK --> PATTERN --> SKILL_REFINE
end
SIG --> CLASS
SKILL_LOOKUP --> EXEC
VERIFY --> OUTCOME
SKILL_REFINE -. informs next decision .-> SKILL_LOOKUP
CLASS -- needs approval --> APPROVAL["Approval request → Matrix ✅/❌"]
APPROVAL -- approved --> EXEC
APPROVAL -- denied --> RESOLVE["Resolve signal (denied)"]
recommended_action lives on the classification, not the signal (SA6): checks
raise facts; the classifier decides what to do about them.
Target architecture
Container stack on mac-mini
One image (oikos:<git-sha>), three long-running roles + init jobs, three Docker
networks as trust boundaries (S10):
net-front— Caddy-facing:apionly.net-data— Postgres + everything that needs it.net-ops— SSH egress: actuator role only (Hermes and API have no SSH).
graph TB
subgraph mac-mini["mac-mini — Docker host (OrbStack), always-on"]
subgraph stack["Docker Compose — image oikos:<sha>"]
MIGRATE["init: oikos migrate + seed\n(one-shot, DDL user)"]
PG["PostgreSQL 16 + TimescaleDB\nentities/relationships • signals\nexecutions • patterns/skills\npolicy • ontology • hypertables"]
INF["Infisical (secrets)"]
API["oikos api\nREST (OpenAPI) + MCP (streamable HTTP)\npolicy enforcement • audit • events\nSSE stream"]
SCHED["oikos scheduler\nchecks (from check_defs) → signals\n+ actuator: classify → execute/escalate\n+ learning engine\nSSH (restricted key) → fleet"]
NOTIF["oikos notifier\nMatrix alerts + approval reactions\nDB rendezvous (no RPC)"]
HERMES["Hermes agent (gateway :8092)\nMCP client → api\nNO SSH keys"]
end
DEPLOY["deploy webhook listener\n(HMAC-verified, non-root)"]
end
subgraph external["External"]
CADDY["Caddy (LXC 121)"]
GITEA["Gitea (LXC 104) + CI"]
MATRIX["Matrix (LXC 118)"]
FLEET["hubris / strong / LXCs / VMs"]
WS["Workstations (Hermes remote)"]
WATCHDOG["Watchdog cron on apps/105\ncurl /healthz → Matrix ping"]
end
MIGRATE --> PG
API --> PG
SCHED --> PG
NOTIF --> PG
HERMES -- "MCP (bearer token)" --> API
SCHED -- "SSH (restricted key)" --> FLEET
NOTIF --- MATRIX
CADDY -- "reverse_proxy + forward-auth" --> API
GITEA -- "webhook (HMAC, CI-gated)" --> DEPLOY
WS -- "mesh, mTLS/token" --> HERMES
WATCHDOG -. "external heartbeat" .-> API
Container hardening: read-only root filesystems, non-root users, restart: always,
stop_grace_period: 30s, images pinned by digest, gcr.io/distroless/static runtime
base, CGO_ENABLED=0, pgx (pure Go).
macOS host realities (R3-13)
Docker on macOS runs in a lightweight VM (use OrbStack: fast, auto-starts at login, stable networking). Consequences the deploy must respect:
- No
network_mode: host. All inbound reachability is via published ports. - Outbound (actuator SSH → hubris/strong/LXCs): direct over the LAN. Containers reach the LAN via Docker NAT with no special config; SSH provides its own encryption, so there is no reason to route this hop through the mesh.
- Inbound: mesh-primary, LAN break-glass. NetBird runs on the macOS host.
- Primary: Caddy and workstations reach the API and Hermes gateway on ports
published on the mesh IP (
<mesh-ip>:8090API,:8092Hermes gateway) — WireGuard-encrypted Caddy→backend hop, mesh membership as a network-level filter, stable addressing, works for roaming workstations. - Break-glass: the API port is also published on the mac-mini's LAN IP
(
<lan-ip>:8090; give mac-mini a DHCP reservation). Safe because the API authenticates every request itself (OIDC JWT / bearer token — S6); the network you arrive from is defense-in-depth, not the auth. This keeps the control plane reachable from inside the house if NetBird's management plane is down after a reboot, and lets the watchdog test the API independently of the mesh. The Hermes gateway stays mesh-only (no LAN binding — it has the weakest application-layer auth story and no break-glass need). - Never bind published ports to
0.0.0.0; enumerate mesh IP, LAN IP (API only), and localhost explicitly.
- Primary: Caddy and workstations reach the API and Hermes gateway on ports
published on the mesh IP (
- Host prep: disable sleep (
sudo pmset -a sleep 0 displaysleep 10), auto-restart after power failure (sudo pmset -a autorestart 1), auto-login enabled so OrbStack starts, macOS auto-updates deferred/scheduled (O7). - Volumes: keep Postgres data on a named Docker volume (VM-native filesystem), not a bind mount — bind mounts cross the VM boundary and are slow.
Multi-node path (designed for, not built)
All roles are stateless; Postgres is the only stateful service. Scaling later means:
run oikos scheduler on another node pointed at the same DB (checks can be
partitioned by a zone attribute on check_defs); run oikos api behind Caddy on
N nodes. SELECT … FOR UPDATE SKIP LOCKED + advisory locks already make the work
queue multi-consumer-safe. Postgres remains the accepted SPOF (mitigated by
backup/DR, below), with an upgrade path to streaming replication if ever needed.
API contract
Contract-first (R3-2)
api/openapi.yaml(OpenAPI 3.1) is the source of truth. CI fails if handlers drift from the spec.- Server:
oapi-codegenstrict-server stubs onchi+ stdlibnet/http. - Clients: the
homelabCLI and future web UIs consume generated clients (Go / TypeScript viaopenapi-typescript). The spec is served atGET /api/v1/openapi.yamland human docs atGET /api/v1/docs(Redoc/Scalar static page) — a UI developer needs nothing but the running API.
Conventions (R3-3)
- Versioning: everything under
/api/v1. Additive-only within v1 (new fields, new endpoints); breaking changes ship as/api/v2side by side with a deprecation window. - Errors: RFC 9457
application/problem+json:{"type":"https://oikos.dev/errors/invalid-transition","title":"invalid lifecycle transition","status":409,"detail":"...","instance":"/api/v1/entities/…","errors":[{field,reason}]}. Domain sentinel errors map centrally:ErrNotFound→404,ErrInvalidTransition→409,ErrApprovalRequired→403,ErrAutonomyBlocked→403,ErrConflict→409,ErrCircuitOpen→503, validation → 422. - Lists: uniform envelope
{"items":[…],"next_cursor":"…"}. Cursor pagination everywhere (?cursor=&limit=, default 50, max 200): keyset on(created_at,id)for entity-ish tables, ontsfor hypertables. MCP tools take the samelimit. - Idempotency: unsafe POSTs (
/executions,/approvals/{id}/decision) accept anIdempotency-Keyheader; keys + response snapshots stored 24h; replay returns the original response. Agents retry safely (R3-3). - Optimistic concurrency: mutable resources carry a
version;GETreturnsETag;PATCH/PUTrequireIf-Match, mismatch → 412. UIs can safely edit. - Timestamps:
TIMESTAMPTZin DB, RFC 3339 UTC on the wire. - CORS: config-driven origin allowlist (empty by default; future UI origins added via config, not code).
- Rate limiting (M2): token bucket per authenticated actor; tighter budget on
/executions; per-(entity,action) cooldown enforced in the actuator besides.
AuthN/AuthZ
| Caller | Mechanism | Scope |
|---|---|---|
| Operator (browser/CLI) | Authentik OIDC; Caddy forward-auth and JWT validated in API middleware (defense in depth, S6/SA10) | role operator (full) or viewer (read-only) |
| Hermes (MCP) | Static bearer token from Infisical, dedicated Docker network, HMAC on requests (S2) | role agent — read tools + POST /executions (which is always policy-gated) |
| Internal roles (scheduler/notifier) | Direct DB with least-privilege DB users; no API hop | n/a |
/healthz, /metrics |
No auth, no audit; not exposed via Caddy (SG18) | n/a |
Roles are claims checked per-route in generated middleware; the spec annotates each operation with its required scope, so future UIs can render capability-aware.
REST surface (summary; the OpenAPI file is normative)
Inventory + ontology:
GET/POST /api/v1/entities,GET/PATCH /api/v1/entities/{id-or-slug}(PATCH covers attribute edits and lifecycle transitions; transition legality validated againstlifecycle_defs)GET /api/v1/entities/{id}/relations,GET /api/v1/graph?root=&depth=&rel_type=→{nodes:[…],edges:[…]}for UI visualization (R3-3)GET /api/v1/ontology(types, relationship types, lifecycles),POST /api/v1/ontology/entity-types,PATCH /api/v1/ontology/entity-types/{name}(policy-gatedconfig_mutation; deprecate-not-delete while instances exist, D3)
Operations:
GET /api/v1/signals,POST /api/v1/signals/{id}/ack|resolve|muteGET /api/v1/checks,POST /api/v1/checks,PATCH /api/v1/checks/{id}(R3-7)GET /api/v1/approvals,POST /api/v1/approvals/{id}/decisionPOST /api/v1/executions(classify → approval check → enqueue; SG15),GET /api/v1/executions/{id},POST /api/v1/executions/{id}/cancelGET /api/v1/classifications?signal_id=…
Learning (operator safety valves, SG7):
GET /api/v1/patterns,PATCH /api/v1/patterns/{id}(activate/invalidate — policy-gatedconfig_mutation, audit-logged)GET /api/v1/skills,GET /api/v1/skills/{id}/versions,PATCH /api/v1/skills/{id}
Policy (dual-control, S3):
GET /api/v1/policy/risk-classes|approval-rules|autonomyPATCHon any policy resource creates a meta-approval; the change applies only after operator approval; before/after hash audit-logged.
Knowledge + observability:
GET /api/v1/knowledge/search?q=,GET /api/v1/knowledge/{entity_id}GET /api/v1/metrics?entity_id=&metric=&from=&to=&rollup=raw|1h|1d(auto-selects resolution by range),GET /api/v1/trends/{entity_id}GET /api/v1/audit?actor=&entity_id=&action=&correlation_id=&from=&to=GET /api/v1/events?type=&entity_id=&severity=&from=&to=andGET /api/v1/events/stream(SSE;Last-Event-IDresume; heartbeat comments; bounded per-subscriber buffers, drop-oldest, P6)GET /api/v1/agent-activity,GET /api/v1/health(fleet summary + trends)GET /api/v1/export— regenerates the three seed YAMLs from DB (round-trip tested, D6)
MCP interface (R3-10)
- Official
github.com/modelcontextprotocol/go-sdk, Streamable HTTP transport, mounted at/mcpon the same binary, bearer-token auth, dedicated network. - Tools delegate to the identical service layer as REST (one behavior, two
protocols):
get_entity,list_entities,get_relations,get_blast_radius,search_knowledge,get_signal_history,get_patterns,get_skills,request_execution,query_metrics,get_trend,get_audit_trail,get_event_timeline,get_agent_activity,get_health_summary. - Docs/runbooks additionally exposed as MCP resources (URI = entity slug) so Hermes can attach them as context without a tool round-trip.
DB-native configuration
seeds/ontology.yaml, seeds/inventory.yaml, seeds/policy.yaml bootstrap the DB
and serve DR; afterwards the DB is authoritative and editable via API. Ingest runs in
the one-shot init container (A4): each seed file applies in a single transaction; a
seed_versions table records applied (file, content-hash) so unchanged seeds are
skipped; GET /api/v1/export regenerates the YAMLs for commit. Round-trip
(seed → DB → export → DB) must be byte-stable — tested in CI.
Config hierarchy (A5): compiled defaults → config file → env vars → Infisical
(secrets only — never plain config). Each role's required keys documented in
docs/operations/config.md.
Repo layout
/
├── docker-compose.yml
├── Makefile # build, test, lint, generate, deploy targets
├── go.mod / go.sum
├── sqlc.yaml
├── .golangci.yml
├── api/
│ └── openapi.yaml # THE API contract (R3-2)
├── cmd/
│ └── oikos/ # single binary: api | scheduler | notifier | all |
│ └── main.go # migrate | seed | export (R3-4)
├── internal/
│ ├── domain/ # pure domain types + state machines + sentinel errors
│ │ ├── entity.go signal.go execution.go classification.go
│ │ ├── pattern.go skill.go approval.go check.go errors.go
│ ├── db/ # pgx pool, sqlc output, repositories (models never escape)
│ │ └── queries/ # sqlc SQL
│ ├── ontology/ # type hierarchy, validation, graph traversal, seed ingest/export
│ ├── httpapi/ # oapi-codegen server impl, middleware (auth, audit,
│ │ # idempotency, rate-limit, problem+json mapping), SSE
│ ├── mcp/ # MCP server (official SDK) over the same services
│ ├── service/ # shared service layer used by httpapi + mcp + loops
│ ├── policy/ # classify, approve (tokens), autonomy, meta-approval
│ ├── scheduler/ # check runner (check_defs → signals/metrics), dedup, flap
│ ├── actuator/ # queue consumer, SSH exec, verify, circuit breaker, locks
│ ├── learning/ # feedback, pattern extraction, skill refinement
│ ├── notifier/ # interface + matrix impl (DB rendezvous)
│ ├── observability/ # slog setup, metrics, audit, events, correlation
│ └── config/
├── migrations/ # golang-migrate, embedded, forward-only (D5/O1)
├── seeds/ # ontology.yaml, inventory.yaml, policy.yaml
├── docs/
│ ├── adr/ # MADR records (R3-16)
│ ├── operations/ # backup-restore.md, dr.md, config.md, runbooks
│ └── … # narrative docs (ingested into knowledge graph)
├── hermes/ # config.yaml, SOUL.md, skills/homelab-ops/SKILL.md
├── compose/
│ ├── oikos/Dockerfile # one multi-stage Dockerfile for the binary
│ └── postgres/ # timescale/timescaledb:2-pg16 config
└── scripts/ # migrate-sops.sh, import-legacy.sh, deploy.sh, watchdog.sh
Developer experience: make generate (oapi-codegen + sqlc), make test
(unit + testcontainers), make dev (docker compose --profile dev up with seeded
fake data + oikos all), .env.example committed.
Database schema
Consolidated migrations — all rev-2 fixes applied inline. Forward-only (no
down.sql; compensating migrations for rollback, plus pre-deploy dumps). Runner:
golang-migrate via oikos migrate in the init container with a DDL-only DB user;
runtime roles get DML-only users (SA9). updated_at maintained by a shared trigger.
001 — Ontology meta-schema (with inheritance, R3-1)
CREATE TABLE lifecycle_defs (
id TEXT PRIMARY KEY, -- 'infrastructure', 'signal', ...
states TEXT[] NOT NULL,
default_state TEXT NOT NULL,
terminal_states TEXT[] NOT NULL DEFAULT '{}',
transitions JSONB NOT NULL, -- {"from":{"to":{"requires":["no-inbound-edges",...]}}}
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE entity_types (
name TEXT PRIMARY KEY, -- 'compute-entity', 'machine', 'proxmox-host'
parent_type TEXT REFERENCES entity_types(name), -- is-a hierarchy (R3-1)
is_abstract BOOLEAN NOT NULL DEFAULT false, -- abstract types can't be instantiated
domain TEXT NOT NULL, -- 'physical','compute','network','storage',
-- 'software','identity','policy','cognition'
layer TEXT NOT NULL CHECK (layer IN ('meta','infrastructure','governance','cognition')),
-- 'meta' is reserved for the abstract root type 'entity'
description TEXT,
lifecycle_id TEXT REFERENCES lifecycle_defs(id),
attribute_schema JSONB, -- JSON Schema for entities.attributes (D2)
schema_version INTEGER NOT NULL DEFAULT 1,
status TEXT NOT NULL DEFAULT 'active', -- 'active'|'deprecated'; no hard delete
-- while instances exist (D3)
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE relationship_types (
name TEXT PRIMARY KEY, -- 'hosts', 'provides', 'depends-on'
inverse TEXT,
source_type TEXT NOT NULL REFERENCES entity_types(name), -- MAY be abstract;
target_type TEXT NOT NULL REFERENCES entity_types(name), -- validation walks hierarchy
cardinality TEXT NOT NULL CHECK (cardinality IN -- source→target multiplicity
('one-to-one','one-to-many','many-to-one','many-to-many')),
description TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE seed_versions ( -- A4: skip unchanged seed files
file TEXT PRIMARY KEY,
content_hash TEXT NOT NULL,
applied_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
Validation semantics (app layer, internal/ontology):
- Instantiating an
is_abstracttype is rejected. - A relationship
(s, t, type)is valid ifftype_of(s)issource_typeor a descendant of it (same for target). The hierarchy is small; resolved in-memory with a cached type tree. - Cardinality enforced by partial unique indexes where expressible
(
one-to-many→UNIQUE (target_id, type);one-to-one→ unique on both ends) plus app-layer checks for the rest.
002 — Entity instances (UUIDv7 + slug, R3-5/D1)
CREATE TABLE entities (
id UUID PRIMARY KEY, -- UUIDv7 generated in Go (time-ordered)
slug TEXT NOT NULL UNIQUE, -- 'host:hubris', 'service:caddy' — human/API handle
type TEXT NOT NULL REFERENCES entity_types(name),
name TEXT NOT NULL,
state TEXT, -- lifecycle state
attributes JSONB NOT NULL DEFAULT '{}', -- validated against attribute_schema
maintenance_until TIMESTAMPTZ, -- R3-8: suppress signals + auto-act while set
version INTEGER NOT NULL DEFAULT 1, -- optimistic lock / ETag source
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (type, name)
);
CREATE INDEX idx_entities_type ON entities(type);
CREATE INDEX idx_entities_state ON entities(state);
CREATE INDEX idx_entities_attrs ON entities USING GIN(attributes);
CREATE TABLE relationships (
source_id UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT,
target_id UUID NOT NULL REFERENCES entities(id) ON DELETE RESTRICT,
type TEXT NOT NULL REFERENCES relationship_types(name),
attributes JSONB,
valid_from TIMESTAMPTZ NOT NULL DEFAULT now(), -- D7: temporal edges
valid_to TIMESTAMPTZ, -- NULL = current
PRIMARY KEY (source_id, target_id, type, valid_from)
);
CREATE INDEX idx_rel_source ON relationships(source_id) WHERE valid_to IS NULL;
CREATE INDEX idx_rel_target ON relationships(target_id) WHERE valid_to IS NULL;
CREATE INDEX idx_rel_type ON relationships(type) WHERE valid_to IS NULL;
-- Cycle-safe traversal (P1): path accumulator prevents revisits; depth capped.
CREATE OR REPLACE FUNCTION blast_radius(start_id UUID, max_depth INT DEFAULT 3,
rel_types TEXT[] DEFAULT NULL)
RETURNS TABLE(entity_id UUID, depth INT) AS $$
WITH RECURSIVE walk AS (
SELECT start_id AS entity_id, 0 AS depth, ARRAY[start_id] AS path
UNION ALL
SELECT r.target_id, w.depth + 1, w.path || r.target_id
FROM relationships r
JOIN walk w ON r.source_id = w.entity_id
WHERE w.depth < LEAST(max_depth, 5)
AND r.valid_to IS NULL
AND NOT r.target_id = ANY(w.path)
AND (rel_types IS NULL OR r.type = ANY(rel_types))
)
SELECT entity_id, MIN(depth) FROM walk GROUP BY entity_id;
$$ LANGUAGE sql STABLE;
Entities are never hard-deleted while edges exist (ON DELETE RESTRICT);
decommission is the lifecycle path (… → destroyed), and destroyed entities remain
as archaeology.
Dual-entity pattern (SA1): every cognition object (signal, classification,
execution, feedback, pattern, skill, approval, check) gets an entities row (so the
graph is traversable: triggers, produces, contributes-to edges live in
relationships) and a typed table below whose PK references entities(id) for
indexed querying.
003 — Operations (signals, checks, approvals, status)
CREATE TABLE check_defs ( -- R3-7: probes as data
entity_id UUID PRIMARY KEY REFERENCES entities(id),
target_id UUID REFERENCES entities(id), -- what it checks (NULL + target_type = type-scoped)
target_type TEXT REFERENCES entity_types(name),
kind TEXT NOT NULL, -- 'http','tcp','disk','cert-expiry','drift','ssh-script'
config JSONB NOT NULL DEFAULT '{}', -- validated per-kind JSON Schema
interval_s INTEGER NOT NULL DEFAULT 600,
timeout_s INTEGER NOT NULL DEFAULT 10,
zone TEXT, -- multi-node partitioning later
enabled BOOLEAN NOT NULL DEFAULT true,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE signals (
entity_id UUID PRIMARY KEY REFERENCES entities(id),
kind TEXT NOT NULL,
severity TEXT NOT NULL CHECK (severity IN ('info','warning','critical')),
target_entity_id UUID REFERENCES entities(id),
check_id UUID REFERENCES check_defs(entity_id),
evidence TEXT,
likely_cause TEXT,
state TEXT NOT NULL DEFAULT 'raised',
occurrence_count INTEGER NOT NULL DEFAULT 1, -- R3-8: dedup counting
first_seen_at TIMESTAMPTZ NOT NULL DEFAULT now(),
last_seen_at TIMESTAMPTZ NOT NULL DEFAULT now(),
flap_count INTEGER NOT NULL DEFAULT 0, -- resolve→re-raise cycles
hold_down_until TIMESTAMPTZ, -- flap suppression window
mute_until TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- R3-8: at most ONE open signal per (target, kind) — repeats update the open row
CREATE UNIQUE INDEX uq_signals_open ON signals(target_entity_id, kind)
WHERE state NOT IN ('resolved','failed');
CREATE INDEX idx_signals_state ON signals(state);
CREATE TABLE approvals (
entity_id UUID PRIMARY KEY REFERENCES entities(id),
subject_entity_id UUID REFERENCES entities(id), -- entity to act on
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
kind TEXT NOT NULL DEFAULT 'execution', -- 'execution'|'policy-change'|'pattern-activation'
payload JSONB, -- e.g. the proposed policy diff
status TEXT NOT NULL DEFAULT 'pending',
token_hash TEXT, -- S5: single-use HMAC token, stored hashed
expires_at TIMESTAMPTZ NOT NULL,
decided_at TIMESTAMPTZ,
decided_by UUID REFERENCES entities(id), -- person entity
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE entity_status ( -- R3-6: replaces state_snapshots (P5)
entity_id UUID PRIMARY KEY REFERENCES entities(id),
health TEXT NOT NULL DEFAULT 'unknown', -- healthy|degraded|down|unknown
last_check_at TIMESTAMPTZ,
details JSONB NOT NULL DEFAULT '{}',
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- health HISTORY is the 'health' metric in metric_samples (retained + rolled up)
CREATE TABLE idempotency_keys ( -- R3-3
key TEXT NOT NULL,
actor TEXT NOT NULL,
request_hash TEXT NOT NULL,
response_code INTEGER,
response_body JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (actor, key)
);
-- pruned by the scheduler after 24h
Approval tokens (S5): token = HMAC(approval_id ‖ subject ‖ action ‖ risk_class ‖ nonce, secret), single-use, stored hashed, TTL-bound; Matrix carries only the approval ID + decision; verification is server-side.
004 — Cognition (classifications, executions, learning)
CREATE TABLE classifications ( -- SA5: every autonomous decision persisted
entity_id UUID PRIMARY KEY REFERENCES entities(id),
signal_entity_id UUID REFERENCES signals(entity_id),
target_entity_id UUID REFERENCES entities(id),
action TEXT NOT NULL,
recommended_action JSONB, -- SA6: lives here, not on the signal
risk_class TEXT NOT NULL,
route TEXT NOT NULL CHECK (route IN ('auto-act','escalate','hold')),
blast_radius UUID[],
pattern_confidence REAL,
skill_id UUID, -- skill entity matched (if any)
autonomy_check TEXT, -- 'allowed' | 'blocked: <reason>'
reasoning JSONB NOT NULL,
correlation_id TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE executions (
entity_id UUID PRIMARY KEY REFERENCES entities(id),
classification_id UUID REFERENCES classifications(entity_id),
signal_entity_id UUID REFERENCES signals(entity_id),
target_entity_id UUID REFERENCES entities(id),
action TEXT NOT NULL,
risk_class TEXT NOT NULL,
approval_id UUID REFERENCES approvals(entity_id),
agent_id UUID REFERENCES entities(id),
skill_id UUID, -- + version pinned at execution time (SG9)
skill_version INTEGER,
status TEXT NOT NULL DEFAULT 'proposed',
result JSONB,
duration_ms INTEGER,
verified BOOLEAN NOT NULL DEFAULT false,
correlation_id TEXT NOT NULL,
started_at TIMESTAMPTZ,
completed_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_exec_target ON executions(target_entity_id);
CREATE INDEX idx_exec_status ON executions(status);
CREATE TABLE feedback (
entity_id UUID PRIMARY KEY REFERENCES entities(id),
execution_id UUID NOT NULL REFERENCES executions(entity_id),
outcome TEXT NOT NULL CHECK (outcome IN ('success','failure','partial','unexpected')),
observation TEXT,
lesson TEXT,
unexpected_side_effects TEXT[],
tags TEXT[],
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_feedback_ts ON feedback(created_at); -- P4: watermark scans
CREATE TABLE patterns (
entity_id UUID PRIMARY KEY REFERENCES entities(id),
applies_type TEXT NOT NULL REFERENCES entity_types(name),
action TEXT NOT NULL,
pattern TEXT NOT NULL,
confidence REAL NOT NULL DEFAULT 0, -- Wilson lower bound, capped by sample size (S4)
evidence_count INTEGER NOT NULL DEFAULT 0,
success_count INTEGER NOT NULL DEFAULT 0,
failure_count INTEGER NOT NULL DEFAULT 0,
status TEXT NOT NULL DEFAULT 'hypothesized',
quarantined BOOLEAN NOT NULL DEFAULT false, -- S4: anomalous feedback bursts
version INTEGER NOT NULL DEFAULT 1, -- D4: optimistic lock
last_validated_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (applies_type, action)
);
-- Counters updated atomically: UPDATE … SET evidence_count = evidence_count + 1 (D4)
CREATE TABLE skills (
entity_id UUID NOT NULL REFERENCES entities(id),
version INTEGER NOT NULL DEFAULT 1, -- SG9: history preserved
name TEXT NOT NULL,
procedure JSONB NOT NULL, -- R3-9: structured, schema-validated (below)
applies_type TEXT REFERENCES entity_types(name),
action TEXT NOT NULL,
pattern_ids UUID[],
status TEXT NOT NULL DEFAULT 'drafted',
success_rate REAL,
changed_by UUID, -- agent or person entity
change_reason TEXT,
last_used_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (entity_id, version)
);
Skill procedure format (R3-9) — validated by JSON Schema at write time; the actuator executes it deterministically; markdown for humans is rendered from it:
{
"params_schema": { "type": "object", "properties": { "unit": {"type":"string"} }, "required": ["unit"] },
"steps": [
{ "name": "restart unit",
"runner": "ssh", "target": "{{ .host }}",
"command": "systemctl restart {{ .unit }}",
"timeout_s": 60 }
],
"verify": [
{ "runner": "ssh", "target": "{{ .host }}",
"command": "systemctl is-active {{ .unit }}",
"expect": { "exit_code": 0, "stdout_contains": "active" },
"retry": { "attempts": 3, "delay_s": 10 } }
],
"rollback": [],
"expected_duration_s": 30,
"known_failure_modes": ["unit masked", "dependency service down"]
}
Templates are Go text/template over validated params; the actuator refuses any
command not generated from a stored skill/runbook step (no free-form agent shell).
005 — Policy
CREATE TABLE risk_classes (
name TEXT PRIMARY KEY, -- read_only, reversible_low, config_mutation, destructive
description TEXT,
approval_required TEXT NOT NULL DEFAULT 'none', -- none|operator|operator_confirmed
autonomy_allowed BOOLEAN NOT NULL DEFAULT false
);
CREATE TABLE approval_rules (
id UUID PRIMARY KEY,
entity_type TEXT REFERENCES entity_types(name), -- may be abstract (R3-1)
action TEXT NOT NULL,
risk_class TEXT NOT NULL REFERENCES risk_classes(name),
autonomy_level TEXT NOT NULL DEFAULT 'escalate' CHECK
(autonomy_level IN ('auto','escalate','never')),
scope_entity UUID REFERENCES entities(id), -- optional per-entity override
version INTEGER NOT NULL DEFAULT 1,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (entity_type, action, scope_entity)
);
CREATE TABLE autonomy_settings (
key TEXT PRIMARY KEY, -- 'global.auto_act', 'never_auto_act.<slug>'
value TEXT NOT NULL,
version INTEGER NOT NULL DEFAULT 1,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
Rule resolution: most-specific wins (scope_entity > concrete type > ancestor type via
the hierarchy). Policy mutations are dual-controlled (S3): the API writes a
policy-change approval; on operator approval the change applies in one transaction
with a before/after hash written to audit_log; at startup each role verifies the
policy hash against last-known-good and raises a critical signal on mismatch.
006 — Observability (TimescaleDB)
Hypertable PKs include the time column (SG1); TimescaleDB DDL is idempotent
(if_not_exists => TRUE, exception-guarded policies — SG3); CAGGs avoid
array_agg (SG2).
CREATE EXTENSION IF NOT EXISTS timescaledb;
CREATE TABLE metric_samples (
ts TIMESTAMPTZ NOT NULL,
entity_id UUID NOT NULL,
metric TEXT NOT NULL, -- 'health','disk_usage_pct','probe_latency_ms',
-- 'api_latency_ms','pattern_confidence',…
value DOUBLE PRECISION NOT NULL,
tags JSONB NOT NULL DEFAULT '{}'
);
SELECT create_hypertable('metric_samples','ts',
chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
CREATE INDEX idx_metrics_entity_ts ON metric_samples(entity_id, ts DESC);
CREATE INDEX idx_metrics_metric_ts ON metric_samples(metric, ts DESC);
DO $$ BEGIN
PERFORM add_retention_policy('metric_samples', INTERVAL '90 days');
EXCEPTION WHEN OTHERS THEN NULL; END $$;
CREATE MATERIALIZED VIEW metric_rollups_1h WITH (timescaledb.continuous) AS
SELECT time_bucket('1 hour', ts) AS bucket, entity_id, metric,
avg(value) AS avg_value, min(value) AS min_value,
max(value) AS max_value, count(*) AS sample_count
FROM metric_samples GROUP BY bucket, entity_id, metric;
-- + metric_rollups_1d identically; refresh policies 1h/1d; rollups kept 1 year
CREATE TABLE audit_log (
id BIGINT GENERATED ALWAYS AS IDENTITY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
actor_type TEXT NOT NULL, -- agent|operator|system|scheduler
actor_id UUID, -- resolves to a real entity (SA3)
action TEXT NOT NULL,
entity_id UUID,
method TEXT, path TEXT, status_code INTEGER,
detail JSONB NOT NULL DEFAULT '{}',
source_ip TEXT,
correlation_id TEXT,
PRIMARY KEY (id, ts) -- SG1
);
SELECT create_hypertable('audit_log','ts',
chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- indexes on (actor_type,actor_id,ts), (entity_id,ts), (correlation_id); 365d retention
CREATE TABLE events (
id BIGINT GENERATED ALWAYS AS IDENTITY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
type TEXT NOT NULL, -- 'signal.raised','execution.completed',…
entity_id UUID,
severity TEXT NOT NULL DEFAULT 'info',
source TEXT NOT NULL,
data JSONB NOT NULL DEFAULT '{}',
correlation_id TEXT,
PRIMARY KEY (id, ts)
);
SELECT create_hypertable('events','ts',
chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- 90d retention; NOTIFY trigger fires after commit for the SSE stream (SG10)
CREATE TABLE agent_activity (
id BIGINT GENERATED ALWAYS AS IDENTITY,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
agent_id UUID NOT NULL,
session_id TEXT,
activity_type TEXT NOT NULL, -- tool_call|reasoning|decision|mcp_query|escalation
tool_name TEXT, entity_id UUID,
input_summary TEXT, output_summary TEXT, -- truncated 500 chars
duration_ms INTEGER, token_count INTEGER, success BOOLEAN,
correlation_id TEXT,
PRIMARY KEY (id, ts)
);
SELECT create_hypertable('agent_activity','ts',
chunk_time_interval => INTERVAL '7 days', if_not_exists => TRUE);
-- 90d retention
-- R3-11: ledger is a VIEW, not a fourth write path
CREATE VIEW ledger AS
SELECT e.created_at AS ts, e.entity_id AS execution_id, e.target_entity_id,
e.action, e.risk_class, e.status, e.verified,
c.route, c.reasoning, a.status AS approval_status, a.decided_by,
e.agent_id, e.correlation_id
FROM executions e
LEFT JOIN classifications c ON c.entity_id = e.classification_id
LEFT JOIN approvals a ON a.entity_id = e.approval_id;
Retention summary: metrics 90d raw / 1y rollups; audit 1y; events 90d; agent
activity 90d; signals/executions/classifications/feedback permanent (they're the
learning corpus and small); idempotency_keys 24h (scheduler prune job).
Control loop
Scheduler (Observe)
- Loads enabled
check_defs; each check runs on its own interval with jitter; a bounded worker pool (errgroup.SetLimit) caps concurrency; per-check timeouts (P2). - Check kinds implemented in Go, config-driven:
http,tcp,disk(SSHdf),cert-expiry,drift(DB inventory vs live state — mismatches raisedriftsignals with the observed diff as evidence),ssh-script(allowlisted). - Every run writes metrics (
health,probe_latency_ms, …) and updatesentity_statusin place (R3-6). - Signal dedup (R3-8): failure upserts against the partial unique index — an
existing open signal gets
occurrence_count+1,last_seen_at=now(). Recovery auto-resolves. A resolve→re-raise cycle incrementsflap_count; after 3 cycles/1h the signal enters hold-down (no notifications, no auto-act) and aflappingmeta-signal is raised for the operator. - Entities with
maintenance_until > now()still get metrics but no signals and are excluded from auto-act. - Housekeeping jobs: daily
pg_dump+ rclone push, idempotency-key prune, feedback watermark advance, secrets-expiry check (M5).
Actuator (Act)
- Consumes signals with
route='auto-act'classifications viaSELECT … FOR UPDATE SKIP LOCKED; per-target serialization withpg_advisory_xact_lock(hashtext(target_entity_id::text))(SG5). - Executes only stored skill/runbook procedures (R3-9) over SSH with a restricted
key: dedicated keypair,
command=/from=constrained inauthorized_keyson targets, mounted read-only into the scheduler/actuator container only (S1). Phase 3 formalizes this as the gateway: all execution flows through/executions; Hermes never touches SSH. - Context-aware SSH (SG13): session started in a goroutine,
ctx.Done()closes session+client to unblock; SSH errors classified (network → retryable/circuit, auth → fatal alert, non-zero exit → failed, timeout →timed_out → verifying). - Circuit breaker per target host (M4): N consecutive failures → open circuit,
exponential backoff, raise
target-unreachablesignal instead of piling failed executions into the learning corpus. - Loop guard: retry budget per (entity, action) from execution history; rate cooldown per (entity, action).
- Autonomy kill-switch consulted every pass (
global.auto_act,never_auto_act.<slug>). - Graceful shutdown (SG4):
signal.NotifyContext; stop intake → 30s drain → if an execution is in flight, markfailed("shutdown interrupted") + emit feedback → close pool. Compose setsstop_grace_period: 30s.
Learning engine
Concrete algorithm (P4 + S4 guardrails):
- Hourly, read feedback past the watermark, joined to executions, grouped by
(applies_type, action). - Update pattern counters atomically; recompute confidence as the Wilson score
lower bound of success rate (conservative for small N), additionally capped by
min(confidence, evidence_count/5)so nothing looks confident before 5 samples. - Status:
hypothesized(N<5) →validated(N≥5 and confidence ≥ 0.7, emitspattern.validatedevent + notification) →activeonly via operator PATCH (policy-gatedconfig_mutation). - Anomaly quarantine: >10 identical-outcome feedback rows within 1h for one
(type, action) →
quarantined=true, meta-signal for review. - Skill refinement: an
activepattern with confidence > 0.7 drafts/refines a skill (new version row;changed_by,change_reasonrecorded). Skills follow their own lifecycle; activation is operator-gated. No skill ever auto-promotes an action intodestructiveautonomy — hard-coded, not policy data. - Classifier consumption: pattern confidence and skill existence feed the auto-act/escalate score; frequent-failure patterns lower confidence.
Notifier
Notifier interface (SendAlert / SendApprovalRequest / decision intake); Matrix
implementation posts approvals with ✅/❌ reactions and writes decisions directly
to the approvals table — the DB is the rendezvous, no service-to-service calls, so
pending approvals survive restarts of either side (SA7/A7). If the notifier is down,
the operator's alternative path is the REST /approvals endpoint.
Security model
Threat model (documented in docs/adr/0007-threat-model.md):
| Boundary | Mechanism |
|---|---|
| Internet/mesh → API | Caddy (TLS) + Authentik forward-auth and in-API OIDC JWT validation (Caddy compromise ≠ API compromise) |
| LAN → API (break-glass) | Same in-API auth (OIDC JWT / bearer) — network origin is defense-in-depth, never the auth. Plaintext hop accepted for emergency/watchdog use only; routine traffic uses the mesh |
| Workstation → Hermes gateway | Mesh membership (network) + gateway token/mTLS (application); mesh-only, no LAN binding |
| Hermes → API (MCP) | Dedicated Docker network + static bearer token + HMAC |
| Roles → Postgres | Least-privilege DB users (DDL only in init; DML per role), TLS on the Docker network (S7) |
| Actuator → fleet | Restricted SSH key (command=,from=), actuator container only; full gateway in Phase 3 (S1/S10) |
| Gitea → deploy | HMAC-signed webhook, localhost/mesh-bound listener, non-root deploy user, CI-gated (S8/M1) |
| Learning → policy | Structurally impossible: learning role's DB user has no write grants on policy tables; changes route through approvals (S3/S4) |
| Approvals | Single-use HMAC tokens, hashed at rest, TTL (S5) |
| Secrets | Infisical with machine identities; bootstrap root of trust = master key in mac-mini Keychain, backed up offline; one SOPS age key retained for DR until a restore drill passes (S9) |
Rotation cadences (M5): SSH actuator key 6mo, Infisical machine tokens 90d, MCP
bearer 90d, webhook HMAC 1y — each with a documented procedure and a scheduler check
that raises a signal 2 weeks before expiry. Supply chain (M6): images pinned by
digest, govulncheck + golangci-lint in CI, distroless runtime.
Observability
All capture goes to Postgres (hypertables above) through internal/observability:
- Metrics — check results, API request count/latency, Go runtime stats, learning
metrics (pattern confidence, skill success rate, auto-act vs escalation ratio),
agent token/tool-call counts. Also exposed as a Prometheus-format
/metricsendpoint (internal only) so Grafana/Prometheus can attach later without schema work. - Audit — middleware on every mutating REST call + MCP tool call + actuator SSH command; policy changes carry before/after hashes; operator identity from OIDC claims (M3).
- Events — emitted in the same transaction as the state change (SG10);
post-commit
NOTIFYfeeds the SSE stream; in-process bus covers API-local events (SG8). SSE subscribers get bounded buffers with drop-oldest + heartbeats (P6); delivery is best-effort, history viaGET /events. - Correlation — a
correlation_idis minted at signal creation (or API request) and propagated viacontext.Contextthrough classification → approval → execution → SSH → verification → feedback, linking audit + events end-to-end. - Logging —
slogJSON to stdout; every line carriesservice,correlation_id,entity_idwhere applicable;debug=trueenables probe payloads/SQL/classification reasoning. - Agent self-inspection — MCP tools let Hermes query its own history, trends, audit trail, and efficiency (token usage over time).
SLOs (M7): check interval 10min ±1min; API p99 < 200ms; deploy < 5min; alert delivery < 30s; watchdog detection < 5min.
Operations
Backup / restore / DR (A3, O3, O4)
- Daily
pg_dump(custom format, compressed) + WAL archiving for PITR; pushed off-host to Proton Drive via rclone (reuse existing rclone credentials from LXC 132 setup). Retention 30 daily + 12 monthly. Pre-deploy dump before every migration run. - Infisical: native backup + secrets exported to one SOPS-age-encrypted file as fallback (the retained age key is the DR escape hatch).
- Monthly automated restore drill: scratch container,
pg_restore, run the export round-trip check, alert on failure. - DR targets: RTO 4h / RPO 24h. Cold-start runbook
(
docs/operations/dr.md): fresh machine → OrbStack + Docker → clone repo → restore Infisical →pg_restore→docker compose up -d→ verify/healthz+ fleet health.
Watchdog (O2)
Cron on apps/105 (outside the stack): every 5min, curl /healthz on both paths —
http://<lan-ip>:8090/healthz (LAN, tests the API itself) and
http://<mesh-ip>:8090/healthz (mesh, tests the path Caddy and workstations use) —
plus pg_isready; on any failure, post directly to the Matrix webhook naming which
path failed (LAN-down = stack problem; mesh-down-LAN-up = NetBird problem). This is
the orthogonal "who watches the watcher" channel.
Deploy, release, rollback (O1, O5, M1, R3-12)
- CI (Gitea Actions):
go vet,golangci-lint run,go test ./... -race -cover(coverage gates: ≥80%internal/policy+internal/learning, ≥60% elsewhere),oapi-codegen/sqlcdiff check (generated code committed and clean),govulncheck,docker build. - Deploy: webhook (HMAC-verified, gated on green CI) →
deploy.sh:git pull→ pre-deploypg_dump→ build imageoikos:<sha>(multi-stage; no hostgo build, P7) →oikos migrateinit job →docker compose up -d --no-depsapp roles (Postgres container never recreated on routine deploys) → healthcheck-gated. - Migration compatibility: additive-only per deploy window (new columns nullable); new code tolerates previous schema for one deploy.
- Rollback: retag compose to the previous SHA (last 5 images kept) → if the migration
was the problem,
pg_restorethe pre-deploy dump →docker compose up -d. Runbook indocs/operations/rollback.md. - Health checks (O6): API
/healthz(DB ping); scheduler freshness ("last successful check pass < 15min") exposed viaentity_statusself-row; notifier "last poll < 60s"; Hermes gateway ping — all wired into composehealthcheck+ watchdog.
Coexistence + cutover (A6)
During coexistence apps/105 is read-only toward shared state (its scheduler disabled once the new stack's checks are verified). Cutover: verify traffic on the new stack → disable apps/105 services → retarget Caddy + Gitea webhooks → cleanup. Rollback after cutover = re-enable apps/105 (its JSONL state frozen at cutover; reconciliation = re-import from Postgres if ever needed).
Phasing — with acceptance criteria (R3-15)
Each phase ends with explicit "done when" checks an implementing agent can run.
Phase 0 — Ontology + contract (no service code):
- Finalize entity types (incl. abstract hierarchy, Person/Agent/IdentityProvider, Cluster/ComposeStack), relationship types (with endpoint types + cardinality), lifecycles (all terminal states + named preconditions).
- Write
seeds/*.yaml; writeapi/openapi.yamlv1 for the full REST surface; write ADRs 0001–0010 from the Decisions table. - Done when: seeds lint against a meta-schema validator script;
openapi.yamlpassesredocly lint; operator has reviewed the lifecycle diagrams + spec.
Phase 1 — Foundation (DB + skeleton):
- Go module
github.com/dtoro/oikos;cmd/oikosskeleton with role subcommands; domain layer + sentinel errors; migrations 001–006;oikos migrate+oikos seed(idempotent, transactional,seed_versions); sqlc + repositories; slog; testcontainers harness; daily backup job + watchdog cron installed. - Done when:
make testgreen including: seed ingest twice = no-op; export round-trip byte-stable;blast_radiuscorrect on a cyclic fixture; abstract-type instantiation rejected; relationship endpoint validation honors inheritance; hypertables + CAGGs + retention created idempotently; legacysignals/*.jsonl+ledger/*.jsonlimported byscripts/import-legacy.sh.
Phase 2 — API:
- oapi-codegen server; auth middleware (OIDC JWT + roles, bearer for agent);
problem+json mapping; idempotency; ETag/If-Match; rate limiting; audit middleware;
transactional event emitter + SSE stream; knowledge ingestion from
docs/(content-hash skip, P3); MCP server (official SDK) over the shared service layer; pattern/skill/policy endpoints with dual-control approvals; CI pipeline live. - Done when: spec-conformance tests pass (schemathesis or generated-client round
trip);
curl /api/v1/entities?type=servicereturns the fleet; MCPlist_entitiesreturns the same data; a mutating call withoutIf-Matchon a stale version → 412; replayedIdempotency-Keyreturns the cached response; SSE stream shows anentity.createdevent; audit rows carry the OIDC sub.
Phase 3 — Control loop:
- Scheduler with
check_defsrunner, dedup/flap/maintenance logic, metrics +entity_status; actuator with restricted SSH key, advisory locks, circuit breaker, retry budgets, graceful shutdown; learning engine per the algorithm above; approval tokens; notifier (Matrix, DB rendezvous). - Done when: killing a probed service raises exactly one signal (repeats
increment
occurrence_count); a flapping fixture enters hold-down; aservice-downsignal on areversible_lowtarget auto-restarts, verifies, and the full correlation chain (signal → classification → execution → SSH audit → feedback) is queryable by onecorrelation_id; adestructiveaction produces a Matrix approval whose ✅ token is single-use; after 5 successful executions a pattern reachesvalidatedand stays there until operator PATCH; kill-switchglobal.auto_act=offforces escalation.
Phase 4 — Agent (Hermes):
- Hermes container (gateway :8092, mesh-published, token/mTLS), homelab skills, MCP wiring, agent-activity logging. No SSH keys in this container.
- Done when: from a workstation over mesh, Hermes answers "what depends on
authentik?" via
get_blast_radius, requests an execution that routes through/executionspolicy gating, and its tool calls appear inagent_activity.
Phase 5 — Secrets (Infisical):
- Infisical up; SOPS migrated; services on machine identities; rotation checks; SOPS-age DR fallback exported.
- Done when: no service reads SOPS at runtime; restore drill of Infisical backup passes; rotation runbooks written.
Phase 6 — Deploy + cutover:
- Full pipeline (CI-gated webhook, SHA images, init migrate, healthcheck rollout);
Caddy re-point (
mcp.hubris.network,oikos.hubris.network→ mac-mini mesh :8090); end-to-end verification below; apps/105 disabled per cutover checklist; first monthly restore drill executed. - Done when: all 14 verification checks pass; watchdog alert fires when the API is stopped manually; rollback drill (previous SHA + pg_restore) rehearsed once.
Verification (end to end)
- Ontology:
SELECT * FROM entity_typesshows the 3-layer hierarchy incl. abstract types; lifecycle defs match the diagrams; abstract instantiation is rejected via API (422). - DB: migrations 001–006 apply idempotently; seeds ingest; re-ingest is a no-op.
- API: REST + MCP return identical data for the fleet; problem+json on errors; pagination envelope everywhere.
- Scheduler: check pass writes metrics +
entity_status; one open signal per (entity, kind) under repeated failure. - Actuator: classify → auto-act or escalate → execute → verify → feedback, all correlation-linked in audit + events.
- Learning: after N=5 similar executions a pattern is
validatedwith a Wilson-bounded confidence; operator PATCH activates it; a skill version appears. - Classifier + autonomy: high-confidence pattern → auto-act on
reversible_low; kill-switch off → always escalate;never_auto_act.<slug>honored. - Hermes: remote workstation session; MCP tools work; activity logged.
- Secrets: services fetch from Infisical; SOPS retired (except DR key).
- Deploy: push → CI → webhook → SHA image → migrate init → rolling restart;
deploy.triggered/deploy.completedevents emitted. - Knowledge:
search_knowledge("caddy")returns docs edged toservice:caddy. - Observability: metrics/trends/audit/health endpoints return correct shapes;
Grafana can read
metric_samplesdirectly. - Correlation tracing: one
correlation_idreconstructs signal → classification → execution → SSH → verification → feedback. - Cutover: apps/105 stopped; watchdog still alive; production traffic served solely by the Docker stack.
Risks / trade-offs
- Go rewrite (~4,400 Python lines replaced) — logic carries over 1:1 (see reuse table); mitigated by the phase gates and the contract-first spec.
- Postgres SPOF — accepted; mitigated by backups/PITR/DR drills; replication is the future path.
- Learning cold start — by design: the agent escalates everything until patterns validate and the operator activates them; trust is earned.
- Single-host mac-mini — watchdog + pmset hardening + documented cold start; multi-node path exists when it matters.
- Infisical bootstrap — SOPS fallback retained until a restore drill passes.
Reuse map (Python → Go)
| Existing | Becomes |
|---|---|
oikos/decide.py |
internal/policy/classify.go (+ pattern/skill inputs) |
oikos/signal.py |
internal/scheduler signal lifecycle (DB-backed, dedup added) |
oikos/approve.py |
internal/policy/approve.go (tokens) + internal/notifier/matrix.go |
oikos/ledger.py |
executions + ledger view + audit_log |
oikos/policy.py + policy.yaml |
internal/policy + seeds/policy.yaml → tables |
oikos/drift.py |
drift check kind |
oikos/relations.py |
internal/ontology/graph.go (SQL traversal) |
oikos/report.py |
report endpoints over DB |
mcp/server.py |
internal/mcp (official SDK) |
bin/homelab |
generated OpenAPI Go client |
oikos/scheduler.py |
internal/scheduler (goroutine worker pool, checks-as-data) |
| (new) | internal/learning, internal/observability, internal/httpapi |
Out of scope (for now)
- Web UI (the OpenAPI contract + SSE +
/graphendpoint are built for it; the UI itself comes later). - Multi-node deployment (designed for; see "Multi-node path").
- Vector embeddings / semantic search (Postgres FTS now;
pgvectoris a schema-only addition later). - WebSocket stream (SSE covers current needs).
- LLM-assisted skill extraction (patterns are statistical for now).
Appendix A — audit resolution ledger
Every rev-2 finding and where rev 3 resolves it. (Full finding text in git history, rev 2 of this file.)
| Finding | Resolution |
|---|---|
| S1, S10 | Security model — restricted SSH key in actuator only; network trust zones; Hermes keyless |
| S2 | MCP bearer token + dedicated network (API contract / Security) |
| S3 | Policy dual-control + hash audit + startup self-check (005 / Security) |
| S4 | Operator-gated pattern activation, Wilson + N/5 cap, quarantine, no destructive auto-promotion (Learning engine) |
| S5 | Single-use HMAC approval tokens, hashed (003) |
| S6, SA10 | In-API OIDC JWT validation; Caddy = explicit trust root (AuthN/AuthZ) |
| S7 | TLS to Postgres; per-role DB users (Security) |
| S8 | HMAC webhook, non-root deploy, CI gate (Deploy) |
| S9 | Infisical bootstrap root of trust + SOPS DR fallback (Security / Phase 5) |
| P1 | Cycle-safe blast_radius with path accumulator + depth cap (002) |
| P2 | Bounded worker pool, jitter, timeouts (Scheduler) |
| P3 | Content-hash skip on knowledge ingestion (Phase 2) |
| P4 | Hourly watermark-based pattern extraction + feedback(created_at) index (004 / Learning) |
| P5 | entity_status replaces state_snapshots; history in retained metrics (R3-6) |
| P6 | SSE bounded buffers, drop-oldest, heartbeats (Observability) |
| P7 | Docker-only builds; no host go build in deploy (Deploy) |
| A1 | Testing strategy embedded in phase gates + CI coverage gates |
| A2 | Observability section + migration 006 |
| A3, O3, O4 | Backup/restore/DR section with drills, RTO/RPO |
| A4 | Init-container migrate/seed, per-file transactions, seed_versions |
| A5 | Config hierarchy (DB-native configuration) |
| A6 | Read-only coexistence + cutover/rollback plan |
| A7, SA7 | Notifier via DB rendezvous |
| D1 | UUIDv7 + slug (R3-5) |
| D2 | attribute_schema JSON Schema validation |
| D3 | entity_type status, deprecate-not-delete, ancestor-aware rules |
| D4 | Atomic counter updates + version optimistic locks |
| D5, O1 | Forward-only migrations + pre-deploy dump + rollback runbook |
| D6 | GET /api/v1/export + byte-stable round-trip test |
| D7 | valid_from/valid_to on relationships |
| O2 | External watchdog cron on apps/105 |
| O5 | --no-deps rollout, pinned Postgres container |
| O6 | Per-role health checks wired to compose + watchdog |
| O7 | pmset hardening, OrbStack autostart, update scheduling (R3-13) |
| M1 | Gitea Actions CI, gated webhook |
| M2 | Per-actor token bucket + /executions budget + actuator cooldowns |
| M3 | audit_log covers operator REST mutations with OIDC identity |
| M4 | Per-target circuit breaker |
| M5 | Rotation cadences + expiry signals |
| M6 | Digest-pinned images, govulncheck, distroless |
| M7 | SLO table (Observability) |
| SA1, SG1 | Dual-entity pattern for all cognition objects; hypertable PKs include ts |
| SA2 | Learning cannot write governance; proposes via approvals (Layer map) |
| SA3 | Person/Agent/IdentityProvider entity types (Governance BDD) |
| SA4 | All lifecycles have terminal states + recovery paths |
| SA5 | classifications table (004) |
| SA6 | recommended_action on classification, not signal |
| SA8 | Cluster, ComposeStack, StandaloneServer attrs (BDD) |
| SA9 | timescale image, init-container migrations, DDL/DML user split |
| SG2, SG3 | CAGGs without array_agg; idempotent TimescaleDB DDL |
| SG4 | Graceful shutdown spec (Actuator) |
| SG5 | Per-entity advisory locks |
| SG6 | internal/domain + repository pattern (Repo layout) |
| SG7 | Pattern/skill PATCH endpoints (API) |
| SG8, SG10 | Transactional events + post-commit NOTIFY + in-process bus |
| SG9 | Skill versions as composite PK + change metadata |
| SG11 | Sentinel errors + problem+json mapping |
| SG13 | Context-aware SSH |
| SG14 | Pool sizing: api 15 / scheduler 5 / actuator 5 / learning 3; max_connections=80; alert at 80% |
| SG15 | POST /api/v1/executions resource style |
| SG16 | Cursor pagination + envelope |
| SG17 | sqlc.yaml, module path, CGO off, pgx, distroless, embedded migrations |
| SG18 | Unauthenticated internal-only /healthz + /metrics |
Appendix B — initial ADRs to write (Phase 0)
0001 Go + single-binary role packaging · 0002 Postgres+TimescaleDB as the only datastore · 0003 DB-native ontology with YAML seeds · 0004 OpenAPI-first API · 0005 UUIDv7 + slug identity · 0006 learning is proposal-only (no self-authorization) · 0007 threat model + trust zones · 0008 forward-only migrations · 0009 SSE over WebSocket · 0010 Infisical with SOPS DR fallback.