From 2d7554436204f52bec365d416a07e1db79ad4426 Mon Sep 17 00:00:00 2001 From: dtoro Date: Mon, 6 Jul 2026 23:06:34 +0200 Subject: [PATCH] plans: SysML BDD ontology, generic model, full audit MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Ontology rewrite: - Replace Mermaid ER diagram with SysML Block Definition Diagrams (BDD) using class diagram syntax: generalization, composition, aggregation, association with multiplicity annotations - Split into 3 diagrams: infrastructure (compute/storage/network), software+services, cognition (operations+learning) - Make compute model generic: ComputeEntity abstract base with specializations (Machine, VirtualMachine, Container→LXC/DockerContainer; Machine→ProxmoxHost/StandaloneServer/Workstation/Appliance) - Hypervisor is software on a Machine (not all machines are Proxmox) - Any ComputeEntity can mount Volumes (VMs AND LXCs, validated) - DockerContainer is first-class (OS models its own infrastructure) - Services on any compute type (not just LXC/VM) - Documents/Runbooks describe any Entity (not just Service) - Added design notes validating assumptions against actual inventory Schema updates: - entity_types: add attribute_schema (JSONB for validating attributes) - entity_types: add status (active/deprecated, no hard delete while instances exist) Audit (37 findings across 6 categories): - Security: 10 findings (5 HIGH) — SSH keys, MCP auth, policy mutability, learning poisoning, confirmation phrase, webhook auth, TLS, blast radius - Performance: 7 findings (1 HIGH) — CTE cycle guard, probe concurrency, ingestion, pattern extraction, table growth, WS backpressure - Architecture: 7 findings (3 HIGH) — testing, observability, DB backup - Data model: 7 findings (1 HIGH) — entity ID, attribute schema, type evolution, concurrent writes, migration rollback, DR export - Operational: 7 findings (4 HIGH) — rollback, watchdog, backup/restore runbook, disaster recovery, deploy downtime, health checks - Missing: 7 findings (1 HIGH) — CI/CD, rate limiting, audit log, circuit breaker, secret rotation, supply chain, SLOs - Top-5 priority items called out before implementation --- ...idate-oikos-control-plane-onto-mac-mini.md | 393 ++++++++++++++++-- 1 file changed, 354 insertions(+), 39 deletions(-) diff --git a/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md b/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md index a4929bc..12f629b 100644 --- a/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md +++ b/plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md @@ -79,13 +79,18 @@ authoritative and editable via API. ```mermaid graph TB - subgraph cognition["Layer 3 — Cognition (the OS's behavior + learning)"] + cognition -- "observes, acts on, learns about" --> infra + governance -- "governs access to" --> infra + governance -- "constrains" --> cognition + cognition -- "creates + refines" --> governance + + subgraph cognition["Layer 3 — Cognition (OS behavior + learning)"] direction LR OBS["Observation\nsignal, state-snapshot"] DEC["Decision\nclassification, risk-assessment"] - ACT["Action\nchange, execution, verification"] + ACT["Action\nexecution, verification"] GOV["Governance\napproval-request, approval-decision"] - KNOW["Knowledge\ndocument, runbook, lesson"] + KNOW["Knowledge\ndocument, runbook"] LEARN["Learning\npattern, skill, feedback"] end @@ -99,52 +104,271 @@ graph TB subgraph infra["Layer 1 — Infrastructure (the managed world)"] direction LR PHYS["Physical\nsite, machine, ups, sensor"] - COMP["Compute\nproxmox-host, lxc, vm,\nworkstation, container"] + COMP["Compute\nmachine, vm, container\n(lxc, docker)"] NET["Network\nlan, mesh, dns-zone,\ningress-route, certificate"] STOR["Storage\nstorage-pool, volume,\nmount, backup-target"] SOFT["Software\nservice, application,\nconfig-repo, deploy-pipeline"] end - - cognition -- "observes, acts on, learns about" --> infra - governance -- "governs access to" --> infra - governance -- "constrains" --> cognition - cognition -- "creates + refines" --> governance ``` -### Entity relationship diagram — core entities and typed edges +### Block definition diagram (SysML BDD) + +Uses Mermaid class diagram syntax following SysML BDD conventions: +- `«abstract»` = abstract block (cannot be instantiated) +- `<\|--` = generalization (is-a) +- `*--` = composition (whole-part, lifecycle dependency) +- `o--` = aggregation (whole-part, independent lifecycle) +- `--` = association (typed link) +- Multiplicity: `"1"`, `"0..1"`, `"1..*"`, `"0..*"`, `"*"` + +**Infrastructure layer — compute, storage, network:** ```mermaid -erDiagram - MACHINE ||--o{ PROXMOX_HOST : "is-a" - PROXMOX_HOST ||--o{ LXC : hosts - PROXMOX_HOST ||--o{ VM : hosts - PROXMOX_HOST ||--o{ WORKSTATION : hosts - LXC ||--o{ SERVICE : provides - VM ||--o{ SERVICE : provides - WORKSTATION ||--o{ SERVICE : provides - LXC ||--o{ MOUNT : has - MOUNT }o--|| STORAGE_POOL : "stores-on" - SERVICE ||--o{ INGRESS_ROUTE : "exposed-by" - INGRESS_ROUTE }o--|| CERTIFICATE : "secured-by" - INGRESS_ROUTE }o--|| IDENTITY_PROVIDER : "secured-by" - SERVICE ||--o{ SERVICE : "depends-on" - CONFIG_REPO ||--|| SERVICE : "configured-by" - DEPLOY_PIPELINE ||--|| SERVICE : "deploys-to" - SERVICE ||--o{ SIGNAL : "monitored-by" - SIGNAL ||--o{ EXECUTION : triggers - EXECUTION ||--|| FEEDBACK : produces - FEEDBACK }o--|| PATTERN : "contributes-to" - PATTERN }o--|| SKILL : "informs" - SKILL ||--o{ CLASSIFICATION : "guides" - CLASSIFICATION ||--|| EXECUTION : "precedes" - EXECUTION ||--o| APPROVAL : "requires" - PERSON ||--o{ APPROVAL : "decides" - DOCUMENT ||--o{ SERVICE : "describes" - RUNBOOK ||--o{ SERVICE : "procedure-for" - AGENT ||--o{ EXECUTION : "performs" - PERSON ||--o{ AGENT : "owns" +classDiagram + class ComputeEntity { + <> + +state lifecycle + +attributes jsonb + } + class Machine { + +cpu_arch + +ram_gb + } + class VirtualMachine { + +vcpus + +memory_mb + +disk_gb + } + class Container { + <> + +runtime + } + class LXC { + +pve_id + +rootfs + } + class DockerContainer { + +image + +compose_stack + } + class ProxmoxHost { + +pve_version + +cluster_member + } + class StandaloneServer { + +hypervisor + } + class Workstation { + +os + +user + } + class Appliance { + +vendor + +model + } + class Hypervisor { + +type + +version + } + + ComputeEntity <|-- Machine + ComputeEntity <|-- VirtualMachine + ComputeEntity <|-- Container + Machine <|-- ProxmoxHost + Machine <|-- StandaloneServer + Machine <|-- Workstation + Machine <|-- Appliance + Container <|-- LXC + Container <|-- DockerContainer + + Machine "1" *-- "0..1" Hypervisor : runs + Hypervisor "1" o-- "0..*" VirtualMachine : hosts + Hypervisor "1" o-- "0..*" Container : hosts + + class StoragePool { + +type lvm, zfs, nfs + +capacity_gb + } + class Volume { + +name + +size_gb + } + class Mount { + +mount_point + +options + } + + StoragePool "1" *-- "0..*" Volume : contains + ComputeEntity "1" o-- "0..*" Mount : has + Mount "0..*" --> "1" Volume : mounts + + class NetworkInterface { + +mac + +ip + } + class Network { + <> + } + class LAN { +subnet} + class Mesh { +provider} + class VLAN { +tag} + + ComputeEntity "1" *-- "0..*" NetworkInterface : has + NetworkInterface "0..*" --> "1" Network : connects-to + Network <|-- LAN + Network <|-- Mesh + Network <|-- VLAN ``` +**Software + services layer:** + +```mermaid +classDiagram + class Service { + +port + +health_url + +risk_notes + } + class Application { + +version + +config + } + class ConfigRepo { + +url + +branch + } + class DeployPipeline { + +trigger + +target_path + } + class IngressRoute { + +pattern + +upstream + } + class Certificate { + +issuer + +expires + } + class DNSZone { + +zone + } + class DNSRecord { + +name + +record_type + +value + } + + ComputeEntity "1" o-- "0..*" Service : provides + Service "1" *-- "0..*" Application : runs + Service "0..1" --> "0..1" ConfigRepo : configured-by + DeployPipeline "0..*" --> "1" Service : deploys-to + IngressRoute "0..*" --> "1" Service : routes-to + IngressRoute "0..*" --> "0..1" Certificate : secured-by + IngressRoute "0..*" --> "0..1" IdentityProvider : secured-by + Service "0..*" --> "0..*" Service : depends-on + DNSZone "1" *-- "0..*" DNSRecord : contains + DNSRecord "0..*" --> "0..1" IngressRoute : resolves-to +``` + +**Cognition layer — operations + learning:** + +```mermaid +classDiagram + class Signal { + +kind + +severity + +state lifecycle + +evidence + +recommended_action + } + class Execution { + +status + +result + +duration_ms + +verified + } + class Feedback { + +outcome + +observation + +lesson + } + class Pattern { + +confidence + +evidence_count + +status + } + class Skill { + +procedure + +version + +success_rate + } + class Classification { + +risk + +route + +reasoning + } + class Approval { + +status + +ttl + } + class Document { + +title + +content + +source_path + } + class Runbook { + +steps + +risk_class + +verification + } + + ComputeEntity "1" o-- "0..*" Signal : monitored-by + Signal "0..1" --> "0..*" Execution : triggers + Classification "1" --> "0..1" Execution : precedes + Execution "1" *-- "0..1" Feedback : produces + Feedback "0..*" --> "0..*" Pattern : contributes-to + Pattern "0..*" --> "0..1" Skill : informs + Skill "0..1" --> "0..*" Classification : guides + Execution "0..1" --> "0..1" Approval : requires + Agent "0..*" --> "0..*" Execution : performs + Person "0..1" --> "0..*" Approval : decides + Person "0..1" --> "0..*" Agent : owns + + Entity "1" o-- "0..*" Document : documented-by + Entity "1" o-- "0..*" Runbook : procedure-for +``` + +### Design notes — validated assumptions + +**Can VMs mount storage pools?** Yes. In Proxmox, VMs have virtual disks on storage +pools (LVM, ZFS, NFS). LXCs have bind mounts and mount points. Both use the same +storage pools. The model reflects this: `ComputeEntity` (abstract) has `Mount` edges +to `Volume`, regardless of whether the compute entity is a VM, LXC, or machine. + +**Not every machine is a Proxmox host.** The current homelab has: 2 Proxmox hosts, 2 +workstations, 1 external VPS, 2 VMs, 19 LXCs. The model uses `Machine` as the base +type with specializations (`ProxmoxHost`, `StandaloneServer`, `Workstation`, +`Appliance`). Only `ProxmoxHost` runs PVE; `StandaloneServer` could run KVM/libvirt; +`Workstation` runs desktop OS + optionally Docker. A `Hypervisor` is software that +runs on a `Machine` and hosts VMs/containers — it's not always present (workstations +and appliances may not have one). + +**Docker containers are first-class compute entities.** The OS itself runs in Docker +containers, and services like Jellyfin's MariaDB sidecar run in Docker within LXCs. +`DockerContainer` is a specialization of `Container` with `image`, `compose_stack` +attributes. This lets the OS model its own infrastructure. + +**Services run on any compute entity.** The homelab has services on LXCs (caddy, +gitea), VMs (zimaos, haos), workstations (mac-mini will run the OS), and external +hosts (authentik on netbird-vps). The `provides` relationship is from +`ComputeEntity` (abstract), not from a specific compute type. + +**Documents and runbooks describe any entity.** The old ER diagram had +`Document → Service` only. In practice, docs describe hosts, containers, network +infrastructure, storage, and investigations. The model uses `Entity` (the root +abstract type) for `documented-by` and `procedure-for`, so any entity can have docs +and runbooks. + ### Infrastructure lifecycle ```mermaid @@ -495,6 +719,8 @@ CREATE TABLE entity_types ( layer TEXT NOT NULL, -- 'infrastructure', 'governance', 'cognition' description TEXT, lifecycle_id TEXT, -- FK to lifecycle_defs (nullable = no lifecycle) + attribute_schema JSONB, -- JSON Schema for validating entity.attributes + status TEXT NOT NULL DEFAULT 'active', -- 'active', 'deprecated' (no hard delete while instances exist) created_at TIMESTAMPTZ DEFAULT now(), updated_at TIMESTAMPTZ DEFAULT now() ); @@ -1033,6 +1259,95 @@ Keep apps/105 running as fallback. Cutover checklist when ready: need to be added/modified. The DB-native approach makes this an API call, not a file edit + redeploy. +## Audit findings (best practices, security, performance, sanity) + +A full audit was performed against the plan. Findings are organized by category and +severity. **High-severity items must be addressed before implementation; medium items +should be addressed during the phase they belong to.** + +### Security + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| S1 | **HIGH** | SSH private keys mounted into Hermes container — a compromised container has unrestricted SSH to all hosts | Use a dedicated, restricted SSH key with `command=` in authorized_keys. Time-box the stopgap. Build the actuator gateway (which brokers SSH per-execution) sooner rather than later. | +| S2 | **HIGH** | MCP interface has no auth — any container on the Docker network has full read access to the control plane | Add a shared secret or mTLS between API and Hermes. Bind MCP to a dedicated Docker network, not the default bridge. Never expose the MCP port via Caddy without auth. | +| S3 | **HIGH** | Policy DB is mutable via the same API it governs — a compromised API can rewrite its own approval rules (e.g. flip `destructive` to auto-act) | Policy mutations require a meta-approval (dual-control). Add an immutable audit log of policy changes. Startup self-check: alert if policy hash differs from a known-good baseline. | +| S4 | **HIGH** | Learning model poisoning — flapping services or misconfigured probes can inject biased feedback to push patterns past the confidence threshold and unlock auto-act for destructive actions | Require human confirmation before pattern transitions to `active`. Cap confidence by sample size (require N≥5 executions). Detect anomalous feedback bursts and quarantine. Never let a skill auto-promote to destructive risk class. | +| S5 | **HIGH** | `confirmation_phrase` is a weak auth primitive — if reused, stored in plaintext, or transmitted via Matrix, it's replayable | Make it single-use (one phrase per approval). Store hashed. Transmit only the decision + HMAC, not the phrase. Prefer signed approval tokens. | +| S6 | MEDIUM | REST `/exec` and `/decide` endpoints rely on Authentik forward-auth only — if the API port is reachable directly on the mesh, auth disappears | Enforce OIDC JWT validation in the API middleware too (defense in depth). Bind API port to localhost + Caddy only. | +| S7 | MEDIUM | No TLS between internal services — Postgres connections are plaintext on the Docker network | TLS to Postgres (server cert verification). Use dedicated Docker networks per trust boundary. Document the threat model: is the Docker network trusted? | +| S8 | MEDIUM | Webhook (Gitea → mac-mini deploy) has no auth specified — anyone who can reach the endpoint can trigger arbitrary code execution via `go build` | HMAC signature verification on the webhook. Bind listener to localhost (Gitea reaches via mesh). Run deploy script as non-root. Verify commit signatures before `git pull`. | +| S9 | MEDIUM | Infisical bootstrapping has a chicken-and-egg — Infisical's own master key must come from somewhere | Document the bootstrap root of trust explicitly: where the master key lives (mac-mini keychain), how it's backed up, revocation path. Keep SOPS as fallback until Infisical has a tested restore-from-backup drill. | +| S10 | MEDIUM | Blast radius of a compromised container is wide — Hermes has SSH + MCP + gateway port + data volume; actuator has SSH + writes to policy/learning tables | Principle of least privilege per container: only the actuator should have SSH, not Hermes. Split networks: data (PG), ops (SSH egress), front (Caddy). Use Docker user namespaces and read-only root filesystems. | + +### Performance + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| P1 | MEDIUM | Recursive CTE `blast_radius` has no cycle guard — cycles (A→B→A) inflate work exponentially with depth | Add a visited-set guard (array accumulator in the recursion). Cap `max_depth` at 3. Add `LIMIT` on the outer query. Consider a closure table for hot-path queries. | +| P2 | MEDIUM | Concurrent probes with no concurrency cap — unbounded goroutines could exhaust FDs or hammer slow targets | Bounded worker pool (`errgroup.SetLimit`). Per-target probe timeout. Jitter to avoid thundering herd on shared backends. | +| P3 | MEDIUM | Knowledge graph re-ingestion on every deploy is wasteful if no docs changed | Content-hash each doc (store hash on the entity). Skip ingestion if hash unchanged. Run in a single transaction with deferred FK checks. | +| P4 | MEDIUM | Pattern extraction frequency unspecified — either too tight (scans whole table) or too loose (patterns lag) | Define cadence (hourly). Use a high-watermark on `feedback.ts`. Add index on `feedback(ts)`. Process only new feedback. | +| P5 | **HIGH** | Table growth unaddressed — `state_snapshots` (every entity every 10 min) and `executions`/`feedback` grow indefinitely | Partition `state_snapshots` and `signals` by month. Define TTLs (snapshots > 90 days → aggregate, raw → archive). Add a `prune_*` job in the scheduler. | +| P6 | MEDIUM | WebSocket `/events` has no backpressure — a stalled client pins a goroutine and accumulates memory | Bounded channel per subscriber with drop-oldest-on-full. Max subscribers cap. Heartbeat/timeout. Document: best-effort vs. guaranteed delivery. | +| P7 | LOW | Deploy script runs `go build` on host AND in Docker — double build, host Go toolchain is a deploy dependency | Standardize on multi-stage Docker build only. Drop host `go build` from deploy script. Keep host toolchain for local dev only. | + +### Architecture + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| A1 | **HIGH** | No testing strategy — 13 workstreams, no mention of unit/integration/property tests | Add a testing workstream: unit tests for `classify.go`, `pattern.go`, `skill.go`; integration tests with testcontainers Postgres; golden-file tests for seed ingest; property test for blast-radius (cycles, depth caps). Coverage gates per package. | +| A2 | **HIGH** | No observability — no structured logging, metrics, or traces. Can't inspect why the classifier escalated or how long probes took | At minimum: structured JSON logs (slog) with correlation IDs per execution. A `/metrics` endpoint (even before full Prometheus). `debug=true` flag for full probe payloads. Reconsider Prometheus deferral — "who watches the watcher" requires metrics. | +| A3 | **HIGH** | DB backup strategy is a one-liner — the DB is the entire control-plane state (ontology, policy, inventory, learning, ledger) | Define: `pg_dump` cadence (daily + WAL archiving for PITR), off-host storage (push to hubris or object storage, not the same mac-mini volume), encryption at rest, tested restore procedure (monthly drill), retention (30 daily + 12 monthly). This is Phase 1, not "once running." | +| A4 | MEDIUM | Seed ingest transactional safety — if API crashes mid-ingest, DB could be in partial state | Run migrations + seed ingest as a distinct init container before the API starts. Wrap each seed file in a single transaction. Add a `seed_version` table to skip already-applied seed versions. | +| A5 | MEDIUM | Config management underspecified — unclear what's env vs. DB vs. file vs. Infisical | Define a config hierarchy: defaults → file → env → Infisical (secrets only). Document each service's required config keys. Avoid putting non-secret config in Infisical. | +| A6 | MEDIUM | apps/105 fallback has no cutover safety — both stacks can't write to the same state without conflict | Define cutover mode: apps/105 is read-only during coexistence. If rollback needed after new OS has been writing to Postgres, document the reconciliation plan. | +| A7 | LOW | Notifier as a separate service is over-engineered initially | Either keep in-process (split later), or document the failure mode: if notifier is down, the API must retry/queue and the operator needs an alternative approval path. | + +### Data model + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| D1 | **HIGH** | `entities.id` as TEXT with manual naming is fragile — renames break the ID; signal IDs require a date-string generator that's race-prone | Use UUID or SERIAL for PK. Keep `name` + `type` as a unique composite for human lookup. Signal IDs: use a DB sequence. | +| D2 | MEDIUM | Ontology flexibility sacrifices type safety — `attributes JSONB` has no schema enforcement per entity type | Add a `schema JSONB` column to `entity_types` (JSON Schema). Validate `attributes` against it on insert/update via a trigger or app-layer check. | +| D3 | MEDIUM | Entity type evolution has no migration story — renaming/removing an entity type cascades through many FK references | Add `status` (`active`/`deprecated`) to `entity_types`. Forbid hard deletes while instances exist. Provide a `merge` endpoint for renames. | +| D4 | MEDIUM | Concurrent writer safety — counters on `patterns` (evidence_count, success_count) have read-modify-write races | Use atomic `UPDATE ... SET evidence_count = evidence_count + 1`. Add `version` optimistic-lock column on `skills` and `patterns`. | +| D5 | MEDIUM | Migration rollback not addressed — only one `.down.sql` shown | Commit to forward-only migrations (compensating migrations for rollbacks) and document it, or write and test every `.down.sql`. | +| D6 | MEDIUM | Data export for DR is asserted but no endpoint or format defined | Add `GET /api/v1/export` (returns the three YAML files regenerated from DB). Test round-trip: seed → DB → export → seed → DB yields identical state. | +| D7 | LOW | `relationships` has no temporal data — can't answer "what depended on authentik last month?" | Add `valid_from`/`valid_to` (nullable = current) for `depends-on` and `hosts` edges. Low priority but cheap to add now. | + +### Operational + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| O1 | **HIGH** | No rollback strategy — if a deploy breaks the API (bad migration, schema bug), recovery is "revert the commit" but the migration may have already run | Define: pre-deploy DB backup snapshot; migration compatibility policy (new code tolerates old schema for one deploy); tested rollback runbook per phase. | +| O2 | **HIGH** | "Who watches the watcher" is unresolved — if the OS is down, no alerts fire. No external/orthogonal monitor for the mac-mini or the OS containers | A minimal external watchdog: a cron on apps/105 (or hubris) that curls the API `/healthz` every 5 min and Matrix-pings the operator directly if it fails. This must be *outside* the Docker stack. | +| O3 | **HIGH** | No backup/restore runbook — "pg_dump" is mentioned but no procedure, tested restore, or definition of what "restored" means | Write `docs/operations/backup-restore.md`: what's backed up (DB, Infisical, Hermes volume, seed YAMLs), where, how often, how to restore each, quarterly restore drill. | +| O4 | **HIGH** | No disaster recovery plan — if the mac-mini dies (disk, theft, water), what's the RTO/RPO? | Define RTO/RPO targets (homelab: RTO 4h, RPO 24h). Name the off-host backup target. Write the DR runbook: fresh mac-mini → install Docker → clone repo → restore Infisical → restore DB → `docker compose up`. | +| O5 | MEDIUM | Deploy downtime — `docker compose up -d` recreates containers; API has a brief gap | Use `docker compose up -d --no-deps` for app services. Don't recreate the Postgres container on routine deploys (pin its image). Healthcheck-gated rollout. | +| O6 | MEDIUM | Health checks asserted but not specified | Define per-service: API `/healthz` (DB ping), scheduler "last successful probe < 15min ago", notifier "last poll < 60s ago", Hermes "gateway responding". Wire into Docker healthcheck + alerting. | +| O7 | MEDIUM | mac-mini single-host failure = total outage — disk failure, macOS update reboot, Docker daemon crash | Automated macOS update deferral/scheduling. Docker `restart: always` on all services. Monitoring heartbeat from external host. Documented cold-start runbook (what comes up first, in what order). | + +### Missing + +| ID | Severity | Finding | Recommendation | +|---|---|---|---| +| M1 | **HIGH** | No CI/CD — deploys are git-push → webhook → build. No linting, no tests before deploy, no gated merges | Add CI stage (Gitea Actions): `go vet`, `golangci-lint`, `go test ./...`, `docker build` (no push). Gate the webhook on green CI. At minimum, deploy script runs `go test` before `docker compose up`. | +| M2 | MEDIUM | No rate limiting on the API — a runaway agent loop or misconfigured skill can hammer the API and DB | Per-caller rate limiting (token bucket) on mutating endpoints. `/exec` needs a per-entity/per-action rate cap to prevent actuator loops. | +| M3 | MEDIUM | No access audit for human activity — the ledger records OS actions, but not operator API actions (entity edits, policy changes, approval decisions) | Add an `audit_log` table for all mutating REST calls, including operator identity from OIDC. | +| M4 | MEDIUM | No circuit breaker for the actuator — if a target host is unreachable, the actuator keeps attempting SSH and generating failed executions, poisoning the learning model | Circuit breaker per target: after N consecutive failures, back off (exponential) and raise a "target unreachable" signal instead of continuing to execute. | +| M5 | MEDIUM | No secret rotation story — SSH keys, Infisical machine tokens, API shared secrets need rotation policies | Define rotation cadences. Document the rotation procedure for each secret class. Add a "secrets expiring" check to the scheduler. | +| M6 | LOW | No dependency/supply-chain hygiene — Go modules, Docker base images, Infisical image not pinned or scanned | Pin base images by digest. Run `govulncheck` in CI. Periodically audit `go.sum`. Use distroless or scratch runtime images. | +| M7 | LOW | No documented SLOs — no quantitative success criteria (probe latency, API p99, deploy time, alert delivery) | Add a small SLO table: probe interval 10min ±1min, API p99 < 200ms, deploy < 5min, alert delivery < 30s. | + +### Top-priority items to address before implementation + +1. **S1 + S10** — SSH keys in containers / wide blast radius → build the actuator gateway first, not incrementally +2. **S3 + S4** — Mutable policy DB + learning model poisoning → these compound: a compromised container can rewrite its own rules and inject feedback to unlock auto-act +3. **A1 + A2 + O2** — No tests, no observability, no external watchdog → can't safely run an autonomous agent without all three +4. **A3 + O3 + O4** — Backup/restore/DR is a one-liner for the SPOF Postgres +5. **O1 + M1** — No rollback strategy and no CI gate on deploys + ## Verification (end to end) 1. **Ontology:** `seeds/ontology.yaml` ingested — `SELECT * FROM entity_types` shows