Problem: runbooks are agent-executable procedures but lived at the repo root, separate from the other agent instruction now under .agents/. Change: - Move runbooks/<name>.md -> .agents/skills/<name>/SKILL.md (folder per skill, matching the wiki-hq skills layout). Frontmatter (name, risk_class, inputs, verification, docs_update_checklist, transition) preserved. - Rewrite links (inbound from plans; between-skill siblings) via the move map. - Update prose references in AGENTS.md, HERMES.md, .agents/OIKOS.md, and the operations schema; fix a pre-existing stale link to operations/commands.md. No code consumed runbooks/ by path, so nothing else changes. Verification: all SKILL.md frontmatter parses with valid risk_class; every lifecycle transition resolves to an oikos/ontology.yaml state; broken-link count 127 -> 126 (fixed one, introduced none). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
3.4 KiB
Oikos metrics stack — Prometheus LXC (planned)
Lifecycle state: planned (see oikos/ontology.yaml
lifecycle). No LXC exists yet — this is the plan doc that state requires
before provisioning starts. Do not add a hosts: entry with a guessed
pve_id until the LXC is actually created; Proxmox assigns the real ID
at pct create time.
Why
Week-3 reliability layer (see OIKOS.md) wants trend signals — "disk full in ~9 days at current rate", temperature creep — which need a real time-series store. The scheduler (oikos/scheduler.py) currently does point-in-time threshold checks only; Prometheus is the one new piece of infrastructure the 30-day roadmap calls for.
Note: undocumented LXC 131 on hubris
Oikos's drift detector (oikos/drift.py) found
pve_id 131 live on hubris (via pct list) with no inventory.yaml
entry — created outside the provision-node runbook, identity unknown
from this repo. Investigate what 131 is before assuming any pve_id is
free; don't let Proxmox auto-assign into a range you haven't confirmed
is actually unused end-to-end.
Plan
- Host: hubris (per the earlier decision: new LXC, not co-located).
- Role:
metrics(Prometheus + local TSDB retention; no Grafana yet — the Week-4 Oikos Console renders its own sparklines from the Prometheus HTTP API, per the plan's Week-3 scope decision). - Networking: LAN + mesh-gated only, no public ingress (matches
homelab_mcp/secrets_issuance'sMESH_SUBNETSpattern) — Prometheus exposes host/service metadata that shouldn't be public. - Scrape targets: node_exporter on hubris and strong (Proxmox hosts)
- any LXC the scheduler needs disk/temp trend data from beyond what
pct/dfalready gives (start with hubris + strong only; expand only if a specific signal needs it).
- any LXC the scheduler needs disk/temp trend data from beyond what
- Storage: local LXC rootfs is sufficient (metrics-only workload,
short retention — no
/mnt/librarymount needed).
Provisioning steps (once pve_id is assigned)
Follow lifecycle-provision-node:
pct create <new-id> ...on hubris — confirm the assigned ID doesn't collide with 131 or anything else live.homelab client add metrics(or the chosen name) withstate: provisioning, stubcontainers/<id>-metrics.md.- Install Prometheus + node_exporter (Debian package or binary release — decide at implementation time; no strong preference recorded here).
- Point node_exporter at hubris + strong (either install locally on each, or scrape via SSH-tunneled metrics — install locally is simpler and is the standard approach).
- Follow lifecycle-activate-node
to flip to
active, complete the doc page, regeneratehosts/*.yaml+infrastructure/topology.md. - Extend
oikos/scheduler.py's disk/temp probes to query Prometheus rate() instead of (or alongside) the current live SSHdfprobe, so drift/threshold Signals gain trend evidence ("full in ~9 days") instead of only a point-in-time percentage.
Open question for the operator
Package choice (apt prometheus vs upstream binary release) and exact
scrape interval aren't decided here — pick at implementation time based
on what's easiest to keep patched via the existing homelab apt-audit/
apt-upgrade fleet tooling if using the Debian package.