18cb79caf9
oikos phase 0: ontology + inventory + policy seeds, OpenAPI contract, ADRs
...
- seeds/ontology.yaml: 59 entity types (5 abstract, is-a hierarchy), 46
relationship types with cardinality, 6 lifecycles with terminal states
and named precondition checks
- seeds/inventory.yaml: 110 entities / 142 relationships translated from
legacy inventory.yaml (fleet, services, ingress, storage, governance,
archaeology); thin spots marked for backfill
- seeds/policy.yaml: 4 risk classes, 27 approval rules (hierarchy-aware,
per-entity overrides), autonomy kill-switch off (cold start)
- api/openapi.yaml: full v1 REST contract (40 paths), RFC 9457 errors,
cursor pagination, idempotency, ETag/If-Match, scopes; redocly-clean
- docs/adr/0001-0010: initial architecture decision records
- scripts/validate-seeds.py: Phase 0 gate — hierarchy, lifecycles,
endpoints, cardinality, policy cross-refs (0 errors)
- plan: layer CHECK gains 'meta' (root type), cardinality gains
'many-to-one'
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 00:17:15 +02:00
ea3b2c3662
plans: oikos rev 3 — consolidated spec, ontology inheritance, OpenAPI-first
...
Merge the rev-2 audit + remediation layers into one self-consistent spec and
close new gaps: meta-schema inheritance (parent_type/is_abstract), contract-
first API (RFC 9457, idempotency, ETag, scopes, /graph), single-binary role
packaging, UUIDv7+slug IDs, checks-as-data, signal dedup/flap/maintenance,
executable skill format, MCP streamable HTTP, SSE events, ledger-as-view,
dual-path networking (mesh-primary + LAN break-glass), per-phase acceptance
criteria, ADRs. Appendix A maps every rev-2 finding to its resolution.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 00:03:25 +02:00
8850b85325
plans: remediate all HIGH audit items + architect/developer review
...
Addresses 15 original HIGH audit findings + 18 new findings from
systems architect + senior Go developer review (572 lines added).
CRITICAL fixes:
- SA1: Cognition objects (execution/feedback/pattern/skill) get dual
entity pattern — entities row + typed table, graph-traversable
- SG1: Hypertable PKs fixed — PRIMARY KEY (id, ts) for audit_log,
events, agent_activity (was id-only, would fail create_hypertable)
HIGH fixes:
- SA2: Remove 'cognition creates governance' arrow (unsupported,
was learning-poisoning vector). Patterns propose, operator accepts.
- SA3: Add Person, Agent, IdentityProvider to ontology (were used
in BDD but never defined)
- SA4: Fix all lifecycle dead-ends — add 'failed' state to infra,
terminal 'failed'/'invalidated' to signals/patterns/skills, add
approval lifecycle diagram, add cancellation/rollback-failure
to executions
- SA5: Add classifications table — persist classifier reasoning
(was modeled in BDD but never stored)
- SA6: Move recommended_action from signals to classifications
- SA8: Add Cluster, ComposeStack, ManagedHost to ontology
- SG2: Drop array_agg from CAGG (unsupported by TimescaleDB)
- SG3: Idempotent TimescaleDB calls (if_not_exists, exception guards)
- SG4: Graceful shutdown (SIGTERM, in-flight protection, 30s grace)
- SG5: Entity-level advisory locks (pg_advisory_xact_lock per target)
- SG6: Domain layer (internal/domain/) — sqlc models never escape db/
Security:
- S1: Restricted SSH key (command=) now + actuator gateway in Phase 3
- S2: MCP shared-secret auth + dedicated Docker network
- S3: Policy mutations require meta-approval (dual-control)
- S4: Pattern activation needs operator confirmation + confidence
capped by sample size (N>=5) + anomaly detection
- S5: Single-use HMAC approval tokens replace confirmation_phrase
- SA10: Gateway mTLS + Caddy as documented trust root + JWT validation
Operational:
- A3/O3/O4: Backup to Proton Drive (daily pg_dump + WAL), restore
runbook, DR plan (RTO 4h, RPO 24h), monthly restore drill
- O1: Forward-only migrations + pre-deploy backup + rollback runbook
- O2: External watchdog cron on apps/105
- M1: CI/CD via Gitea Actions (go vet, lint, test -race, docker build)
Architecture:
- A1: Testing strategy with specific tests per package + coverage gates
- SA7: Notifier decoupled via DB rendezvous (no service-to-service calls)
- SA9: TimescaleDB Docker image specified + init container for migrations
- SG7: Pattern/skill management endpoints (operator override)
- SG8: WebSocket push via in-process bus + LISTEN/NOTIFY
- SG10: Transactional event emission (same tx as state change)
- SG11: Error handling — sentinel errors + HTTP mapping + SSH taxonomy
- SG13: Context-aware SSH (x/crypto/ssh doesn't honor context)
- SG14: Connection pool sizing (28 total, max_connections=80)
- SG15: RESTful /executions (was /exec)
- SG16: Pagination on all list endpoints
- SG17: Go tooling (sqlc.yaml, module path, CGO_ENABLED=0, distroless)
- SG18: /healthz and /metrics bypass auth + audit
Updated phasing incorporates all remediation.
2026-07-06 23:35:50 +02:00
3a35289f46
plans: add observability — metrics, audit log, events, agent activity
...
Adds comprehensive data capture layer to the OS plan:
1. TimescaleDB — PostgreSQL extension for time-series data. No separate
database. Hypertables auto-partition, continuous aggregates provide
1h/1d rollups, retention policies auto-drop old data.
2. Migration 6 — 4 new hypertables:
- metric_samples: generic time-series (health, disk, latency, API p99,
goroutines, pattern_confidence, skill_success_rate, agent tokens)
- audit_log: immutable who-did-what trail (every mutating API call,
MCP tool call, SSH command, policy change) — 1 year retention
- events: structured state-change feed (signal lifecycle, execution
lifecycle, approval, deploy, learning, entity, policy) — 90 days
- agent_activity: Hermes tool calls, reasoning, token usage, latency
— 90 days
3. Correlation IDs — propagated through the full call chain (signal →
classification → execution → SSH → verification → feedback → pattern)
so any action chain can be reconstructed end-to-end.
4. 6 new MCP tools for agent self-query:
query_metrics, get_trend, get_audit_trail, get_event_timeline,
get_agent_activity, get_health_summary
5. 8 new REST endpoints for metrics/audit/events/health/trends/export
6. Workstream 14 — observability + data capture (7 sub-components A-G):
metrics, audit, events, agent activity, structured logging, agent
data availability, future visualization plug-in points
7. New Mermaid diagram — observability data capture and query flow
8. Updated architecture diagram to show observability data flows
9. Updated phasing — observability woven into phases 1-3
10. Updated verification — 14 end-to-end checks (was 12), including
metrics querying, audit trail, correlation tracing
2026-07-06 23:12:34 +02:00
2d75544362
plans: SysML BDD ontology, generic model, full audit
...
Ontology rewrite:
- Replace Mermaid ER diagram with SysML Block Definition Diagrams (BDD)
using class diagram syntax: generalization, composition, aggregation,
association with multiplicity annotations
- Split into 3 diagrams: infrastructure (compute/storage/network),
software+services, cognition (operations+learning)
- Make compute model generic: ComputeEntity abstract base with
specializations (Machine, VirtualMachine, Container→LXC/DockerContainer;
Machine→ProxmoxHost/StandaloneServer/Workstation/Appliance)
- Hypervisor is software on a Machine (not all machines are Proxmox)
- Any ComputeEntity can mount Volumes (VMs AND LXCs, validated)
- DockerContainer is first-class (OS models its own infrastructure)
- Services on any compute type (not just LXC/VM)
- Documents/Runbooks describe any Entity (not just Service)
- Added design notes validating assumptions against actual inventory
Schema updates:
- entity_types: add attribute_schema (JSONB for validating attributes)
- entity_types: add status (active/deprecated, no hard delete while instances exist)
Audit (37 findings across 6 categories):
- Security: 10 findings (5 HIGH) — SSH keys, MCP auth, policy mutability,
learning poisoning, confirmation phrase, webhook auth, TLS, blast radius
- Performance: 7 findings (1 HIGH) — CTE cycle guard, probe concurrency,
ingestion, pattern extraction, table growth, WS backpressure
- Architecture: 7 findings (3 HIGH) — testing, observability, DB backup
- Data model: 7 findings (1 HIGH) — entity ID, attribute schema, type
evolution, concurrent writes, migration rollback, DR export
- Operational: 7 findings (4 HIGH) — rollback, watchdog, backup/restore
runbook, disaster recovery, deploy downtime, health checks
- Missing: 7 findings (1 HIGH) — CI/CD, rate limiting, audit log, circuit
breaker, secret rotation, supply chain, SLOs
- Top-5 priority items called out before implementation
2026-07-06 23:06:34 +02:00
d44979aca7
plans: rev 2 — Go rewrite, ontology-first, DB-native config, learning loop
...
Major revision of the Docker-based homelab OS plan:
1. Go instead of Python — all services rewritten as Go binaries
(Gin web framework, sqlc for DB access, goroutines for probes)
2. Ontology-first design — systems modeling with 3 layers:
- Infrastructure (physical, compute, network, storage, software)
- Governance (identity, secrets, policy)
- Cognition (observation, decision, action, knowledge, learning)
7 Mermaid diagrams: layer map, ER diagram, 3 lifecycle state machines,
feedback loop, policy model
3. DB-native config — inventory.yaml/ontology.yaml/policy.yaml become
seed manifests (bootstrap + DR). The DB is the runtime source of truth,
editable via API. Ontology IS the DB schema (entity_types,
relationship_types, lifecycle_defs tables).
4. Feedback loop — agent learns from execution:
execution → outcome → feedback → pattern → skill → classification
Patterns accumulate from execution history, skills codify proven
procedures, classifier uses pattern confidence for auto-act decisions.
Cold start: agent starts cautious, earns autonomy through evidence.
5. 5 migration groups: ontology meta-schema, entity instances, operations,
learning model, policy. Recursive blast_radius SQL function.
6. Phase 0 added: ontology design before any code.
2026-07-06 22:50:49 +02:00
fe54af30f6
plans: replace ASCII architecture diagrams with Mermaid
...
4 diagrams: container stack, OODA loop, knowledge graph, deploy flow
2026-07-06 22:34:53 +02:00
bead722fac
plans: pivot oikos consolidation to docker-based agentic homelab OS
...
Supersedes the launchd-based consolidation plan. Key changes:
- Docker-based deployment on mac-mini (docker compose)
- PostgreSQL for all mutable state (signals, ledger, knowledge graph)
- Infisical replaces SOPS+age for secrets management
- Unified API merges MCP server + homelab CLI (REST + MCP interfaces)
- Hermes agent runs in Docker (gateway mode, connect from any workstation)
- Knowledge graph in Postgres replaces narrative wiki files as agent context
- Structured entity relationships link docs to inventory entities
- Hybrid SSH access (mounted keys now, actuator gateway later)
- Git push → Gitea webhook → Docker rebuild = deploy trigger
- 6-phase rollout: DB → services → agent → secrets → deploy → cutover
2026-07-06 22:32:21 +02:00
14448a7dd9
chore: add plan
2026-07-06 21:56:56 +02:00