Architecture has changed drastically (hexagonal refactor, web client extraction). Every active plan has been reviewed, annotated with 'Completed' or 'Won't do' status, and moved to plans/done/. Completed (7): gaps-and-improvements, liveness-drift, gated-execution, nomos-code-review, codebase-cleanup, mascot-physics, backend-eval Won't do (7): prometheus-lxc, control-room-webui, activity-gaps, activity-timeline, frontend-os-apps, haos-capability-gaps, arr-audit
3.7 KiB
Oikos metrics stack — Prometheus LXC (planned)
Reviewed 2026-08-16 — Status: Won't do — planned but never started; file paths and architecture references are stale post-hexagonal refactor.
Lifecycle state: planned (see seeds/ontology.yaml
lifecycle). No LXC exists yet — this is the plan doc that state requires
before provisioning starts. Do not add a host entry with a guessed
pve_id until the LXC is actually created; Proxmox assigns the real ID
at pct create time.
Why
Week-3 reliability layer (see OIKOS.md) wants trend signals — "disk full in ~9 days at current rate", temperature creep — which need a real time-series store. The Go scheduler (internal/scheduler/scheduler.go) currently does point-in-time threshold checks only; Prometheus is the one new piece of infrastructure the 30-day roadmap calls for.
Note: LXC 131 is teddycloud
The drift detector (drift check kind in check_defs) found
pve_id 131 live on hubris — this is teddycloud (lxc:teddycloud),
which is now documented in seeds/knowledge.yaml. The pve_id range is
fully accounted for (101-134). Next available VMIDs start at 135.
Plan
- Host: hubris (per the earlier decision: new LXC, not co-located).
- Role:
metrics(Prometheus + local TSDB retention; no Grafana yet — the Week-4 Oikos Console renders its own sparklines from the Prometheus HTTP API, per the plan's Week-3 scope decision). - Networking: LAN + mesh-gated only, no public ingress (matches
mcp/secrets-issuanceMESH_SUBNETSpattern) — Prometheus exposes host/service metadata that shouldn't be public. - Scrape targets: node_exporter on hubris and strong (Proxmox hosts)
- any LXC the scheduler needs disk/temp trend data from beyond what
pct/dfalready gives (start with hubris + strong only; expand only if a specific signal needs it).
- any LXC the scheduler needs disk/temp trend data from beyond what
- Storage: local LXC rootfs is sufficient (metrics-only workload,
short retention — no
/mnt/librarymount needed).
Provisioning steps (once pve_id is assigned)
- Add entity to
seeds/inventory.yamlwith stateplannedand the assigned pve_id. Runoikos seedto ingest. pct create <new-id> ...on hubris — next available PVE ID is 135.- Install Prometheus + node_exporter (Debian package or binary release — decide at implementation time; no strong preference recorded here).
- Point node_exporter at hubris + strong (either install locally on each, or scrape via SSH-tunneled metrics — install locally is simpler).
- Transition to
activeviaPATCH /api/v1/entities/{slug}or operator approval through the execution flow. - Add a
prometheuscheck kind tocheck_defsin the scheduler (internal/scheduler/scheduler.go) so trend signals can query Prometheusrate()alongside the current SSHdfprobe, giving "full in ~9 days" predictions instead of only point-in-time percentages. Register aprometheuscheck def viaPOST /api/v1/checks. - Run
oikos exportto regenerateseeds/inventory.yamlfor git.
Open question for the operator
Package choice (apt prometheus vs upstream binary release) and exact
scrape interval aren't decided here — pick at implementation time based
on what's easiest to keep patched via MCP request_execution(action="apt_upgrade").
Changelog
2026-07-08 — Go references updated
Replaced Python references (oikos/scheduler.py, oikos/drift.py,
bin/homelab, homelab client add) with Go equivalents: internal/scheduler/,
check_defs drift kind, MCP request_execution, seeds/inventory.yaml.
LXC 131 identified as teddycloud (PVE range 101-134 fully accounted for).