# Oikos metrics stack — Prometheus LXC (planned) Lifecycle state: **planned** (see [oikos/ontology.yaml](../oikos/ontology.yaml) lifecycle). No LXC exists yet — this is the plan doc that state requires before provisioning starts. Do not add a `hosts:` entry with a guessed `pve_id` until the LXC is actually created; Proxmox assigns the real ID at `pct create` time. ## Why Week-3 reliability layer (see [OIKOS.md](../.agents/OIKOS.md)) wants trend signals — "disk full in ~9 days at current rate", temperature creep — which need a real time-series store. The scheduler ([oikos/scheduler.py](../oikos/scheduler.py)) currently does point-in-time threshold checks only; Prometheus is the one new piece of infrastructure the 30-day roadmap calls for. ## Note: undocumented LXC 131 on hubris Oikos's drift detector ([oikos/drift.py](../oikos/drift.py)) found `pve_id 131` live on hubris (via `pct list`) with no `inventory.yaml` entry — created outside the provision-node runbook, identity unknown from this repo. **Investigate what 131 is before assuming any pve_id is free**; don't let Proxmox auto-assign into a range you haven't confirmed is actually unused end-to-end. ## Plan - **Host:** hubris (per the earlier decision: new LXC, not co-located). - **Role:** `metrics` (Prometheus + local TSDB retention; no Grafana yet — the Week-4 Oikos Console renders its own sparklines from the Prometheus HTTP API, per the plan's Week-3 scope decision). - **Networking:** LAN + mesh-gated only, no public ingress (matches `homelab_mcp`/`secrets_issuance`'s `MESH_SUBNETS` pattern) — Prometheus exposes host/service metadata that shouldn't be public. - **Scrape targets:** node_exporter on hubris and strong (Proxmox hosts) + any LXC the scheduler needs disk/temp trend data from beyond what `pct`/`df` already gives (start with hubris + strong only; expand only if a specific signal needs it). - **Storage:** local LXC rootfs is sufficient (metrics-only workload, short retention — no `/mnt/library` mount needed). ## Provisioning steps (once pve_id is assigned) Follow [lifecycle-provision-node](../.agents/skills/lifecycle-provision-node/SKILL.md): 1. `pct create ...` on hubris — confirm the assigned ID doesn't collide with 131 or anything else live. 2. `homelab client add metrics` (or the chosen name) with `state: provisioning`, stub `containers/-metrics.md`. 3. Install Prometheus + node_exporter (Debian package or binary release — decide at implementation time; no strong preference recorded here). 4. Point node_exporter at hubris + strong (either install locally on each, or scrape via SSH-tunneled metrics — install locally is simpler and is the standard approach). 5. Follow [lifecycle-activate-node](../.agents/skills/lifecycle-activate-node/SKILL.md) to flip to `active`, complete the doc page, regenerate `hosts/*.yaml` + `infrastructure/topology.md`. 6. Extend `oikos/scheduler.py`'s disk/temp probes to query Prometheus rate() instead of (or alongside) the current live SSH `df` probe, so drift/threshold Signals gain trend evidence ("full in ~9 days") instead of only a point-in-time percentage. ## Open question for the operator Package choice (apt `prometheus` vs upstream binary release) and exact scrape interval aren't decided here — pick at implementation time based on what's easiest to keep patched via the existing `homelab apt-audit`/ `apt-upgrade` fleet tooling if using the Debian package.