Files
oikos/plans/2026-07-05-oikos-prometheus-lxc.md
dtoro b5c1247093 docs: streamline & consolidate the tree (phase 6)
Problem: after the wiki-hq reorg, agent-instruction and human-doc domains
were still scattered across the repo root, with three now-redundant stub
files cluttering it. The organizing principle wasn't visible in the layout.

Change — enforce three clear buckets:
- .agents/  = how agents operate: OIKOS.md, HERMES.md (moved from root),
  shared/ conventions, domains/ schemas, skills/, and operations/ (operator
  cheatsheet + enrollment + hermes-agent, moved from root).
- knowledge/ = what exists + evidence: wiki/, GLOSSARY.md, and sources/ now
  including investigations/ (incident records are evidence/sources).
- root = substrate + two entry points (AGENTS.md, README.md), plus plans/
  as its own design-intent domain.

Moves:
- investigations/ -> knowledge/sources/investigations/ (incl. archive/, index).
- operations/ -> .agents/operations/.
- HERMES.md -> .agents/HERMES.md.
- Deleted unreferenced root stubs CAVEMAN.md, CONTRIBUTING.md, and OIKOS.md
  (its 7 remaining linkers repointed to .agents/OIKOS.md).

Consumers updated:
- inventory.yaml doc_page (agent-enrollment) + regenerated hosts/*.yaml + cards.
- tools/setup-hermes-soul.sh and bootstrap.sh (x2) -> .agents/HERMES.md.
- bin/homelab help string -> .agents/operations/hermes-agent.md.
- knowledge/operations schemas, llm-wiki, page-templates, incident-investigation
  skill, AGENTS.md/README nav -> new investigations/operations paths.
- All markdown links rewritten via the path-resolving mapper.

Left in place (substrate/executable/separate-domain): hosts/, ledger/, tools/,
plans/, oikos/, mcp/, secrets/, bin/, inventory.yaml.

Verification: docs-lint at baseline (2 intentional cross-repo refs, no new
breakage); gen-topology.py --check exit 0; build_host_files.py idempotent; all
doc_page targets resolve; Hermes provisioning scripts point at the new path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 18:12:14 +02:00

70 lines
3.4 KiB
Markdown

# Oikos metrics stack — Prometheus LXC (planned)
Lifecycle state: **planned** (see [oikos/ontology.yaml](../oikos/ontology.yaml)
lifecycle). No LXC exists yet — this is the plan doc that state requires
before provisioning starts. Do not add a `hosts:` entry with a guessed
`pve_id` until the LXC is actually created; Proxmox assigns the real ID
at `pct create` time.
## Why
Week-3 reliability layer (see [OIKOS.md](../.agents/OIKOS.md)) wants trend
signals — "disk full in ~9 days at current rate", temperature creep —
which need a real time-series store. The scheduler
([oikos/scheduler.py](../oikos/scheduler.py)) currently does point-in-time
threshold checks only; Prometheus is the one new piece of infrastructure
the 30-day roadmap calls for.
## Note: undocumented LXC 131 on hubris
Oikos's drift detector ([oikos/drift.py](../oikos/drift.py)) found
`pve_id 131` live on hubris (via `pct list`) with no `inventory.yaml`
entry — created outside the provision-node runbook, identity unknown
from this repo. **Investigate what 131 is before assuming any pve_id is
free**; don't let Proxmox auto-assign into a range you haven't confirmed
is actually unused end-to-end.
## Plan
- **Host:** hubris (per the earlier decision: new LXC, not co-located).
- **Role:** `metrics` (Prometheus + local TSDB retention; no Grafana yet —
the Week-4 Oikos Console renders its own sparklines from the Prometheus
HTTP API, per the plan's Week-3 scope decision).
- **Networking:** LAN + mesh-gated only, no public ingress (matches
`homelab_mcp`/`secrets_issuance`'s `MESH_SUBNETS` pattern) — Prometheus
exposes host/service metadata that shouldn't be public.
- **Scrape targets:** node_exporter on hubris and strong (Proxmox hosts)
+ any LXC the scheduler needs disk/temp trend data from beyond what
`pct`/`df` already gives (start with hubris + strong only; expand only
if a specific signal needs it).
- **Storage:** local LXC rootfs is sufficient (metrics-only workload,
short retention — no `/mnt/library` mount needed).
## Provisioning steps (once pve_id is assigned)
Follow [lifecycle-provision-node](../.agents/skills/lifecycle-provision-node/SKILL.md):
1. `pct create <new-id> ...` on hubris — confirm the assigned ID doesn't
collide with 131 or anything else live.
2. `homelab client add metrics` (or the chosen name) with `state:
provisioning`, stub `containers/<id>-metrics.md`.
3. Install Prometheus + node_exporter (Debian package or binary release —
decide at implementation time; no strong preference recorded here).
4. Point node_exporter at hubris + strong (either install locally on each,
or scrape via SSH-tunneled metrics — install locally is simpler and is
the standard approach).
5. Follow [lifecycle-activate-node](../.agents/skills/lifecycle-activate-node/SKILL.md)
to flip to `active`, complete the doc page, regenerate
`hosts/*.yaml` + `infrastructure/topology.md`.
6. Extend `oikos/scheduler.py`'s disk/temp probes to query Prometheus
rate() instead of (or alongside) the current live SSH `df` probe, so
drift/threshold Signals gain trend evidence ("full in ~9 days") instead
of only a point-in-time percentage.
## Open question for the operator
Package choice (apt `prometheus` vs upstream binary release) and exact
scrape interval aren't decided here — pick at implementation time based
on what's easiest to keep patched via the existing `homelab apt-audit`/
`apt-upgrade` fleet tooling if using the Debian package.