Problem: runbooks are agent-executable procedures but lived at the repo root, separate from the other agent instruction now under .agents/. Change: - Move runbooks/<name>.md -> .agents/skills/<name>/SKILL.md (folder per skill, matching the wiki-hq skills layout). Frontmatter (name, risk_class, inputs, verification, docs_update_checklist, transition) preserved. - Rewrite links (inbound from plans; between-skill siblings) via the move map. - Update prose references in AGENTS.md, HERMES.md, .agents/OIKOS.md, and the operations schema; fix a pre-existing stale link to operations/commands.md. No code consumed runbooks/ by path, so nothing else changes. Verification: all SKILL.md frontmatter parses with valid risk_class; every lifecycle transition resolves to an oikos/ontology.yaml state; broken-link count 127 -> 126 (fixed one, introduced none). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
70 lines
3.4 KiB
Markdown
70 lines
3.4 KiB
Markdown
# Oikos metrics stack — Prometheus LXC (planned)
|
|
|
|
Lifecycle state: **planned** (see [oikos/ontology.yaml](../oikos/ontology.yaml)
|
|
lifecycle). No LXC exists yet — this is the plan doc that state requires
|
|
before provisioning starts. Do not add a `hosts:` entry with a guessed
|
|
`pve_id` until the LXC is actually created; Proxmox assigns the real ID
|
|
at `pct create` time.
|
|
|
|
## Why
|
|
|
|
Week-3 reliability layer (see [OIKOS.md](../OIKOS.md)) wants trend
|
|
signals — "disk full in ~9 days at current rate", temperature creep —
|
|
which need a real time-series store. The scheduler
|
|
([oikos/scheduler.py](../oikos/scheduler.py)) currently does point-in-time
|
|
threshold checks only; Prometheus is the one new piece of infrastructure
|
|
the 30-day roadmap calls for.
|
|
|
|
## Note: undocumented LXC 131 on hubris
|
|
|
|
Oikos's drift detector ([oikos/drift.py](../oikos/drift.py)) found
|
|
`pve_id 131` live on hubris (via `pct list`) with no `inventory.yaml`
|
|
entry — created outside the provision-node runbook, identity unknown
|
|
from this repo. **Investigate what 131 is before assuming any pve_id is
|
|
free**; don't let Proxmox auto-assign into a range you haven't confirmed
|
|
is actually unused end-to-end.
|
|
|
|
## Plan
|
|
|
|
- **Host:** hubris (per the earlier decision: new LXC, not co-located).
|
|
- **Role:** `metrics` (Prometheus + local TSDB retention; no Grafana yet —
|
|
the Week-4 Oikos Console renders its own sparklines from the Prometheus
|
|
HTTP API, per the plan's Week-3 scope decision).
|
|
- **Networking:** LAN + mesh-gated only, no public ingress (matches
|
|
`homelab_mcp`/`secrets_issuance`'s `MESH_SUBNETS` pattern) — Prometheus
|
|
exposes host/service metadata that shouldn't be public.
|
|
- **Scrape targets:** node_exporter on hubris and strong (Proxmox hosts)
|
|
+ any LXC the scheduler needs disk/temp trend data from beyond what
|
|
`pct`/`df` already gives (start with hubris + strong only; expand only
|
|
if a specific signal needs it).
|
|
- **Storage:** local LXC rootfs is sufficient (metrics-only workload,
|
|
short retention — no `/mnt/library` mount needed).
|
|
|
|
## Provisioning steps (once pve_id is assigned)
|
|
|
|
Follow [lifecycle-provision-node](../.agents/skills/lifecycle-provision-node/SKILL.md):
|
|
|
|
1. `pct create <new-id> ...` on hubris — confirm the assigned ID doesn't
|
|
collide with 131 or anything else live.
|
|
2. `homelab client add metrics` (or the chosen name) with `state:
|
|
provisioning`, stub `containers/<id>-metrics.md`.
|
|
3. Install Prometheus + node_exporter (Debian package or binary release —
|
|
decide at implementation time; no strong preference recorded here).
|
|
4. Point node_exporter at hubris + strong (either install locally on each,
|
|
or scrape via SSH-tunneled metrics — install locally is simpler and is
|
|
the standard approach).
|
|
5. Follow [lifecycle-activate-node](../.agents/skills/lifecycle-activate-node/SKILL.md)
|
|
to flip to `active`, complete the doc page, regenerate
|
|
`hosts/*.yaml` + `infrastructure/topology.md`.
|
|
6. Extend `oikos/scheduler.py`'s disk/temp probes to query Prometheus
|
|
rate() instead of (or alongside) the current live SSH `df` probe, so
|
|
drift/threshold Signals gain trend evidence ("full in ~9 days") instead
|
|
of only a point-in-time percentage.
|
|
|
|
## Open question for the operator
|
|
|
|
Package choice (apt `prometheus` vs upstream binary release) and exact
|
|
scrape interval aren't decided here — pick at implementation time based
|
|
on what's easiest to keep patched via the existing `homelab apt-audit`/
|
|
`apt-upgrade` fleet tooling if using the Debian package.
|