complete comprehensive audit — all cleanup items resolved

Plan #4 done. Audit inventory verified against codebase:

- 9 superseded oikos/*.py files deleted (only gen-topology.py remains)
- bin/homelab deleted, bin/oikos deleted, oikos/cards/ deleted
- .hermes/plans/ already archived to archive/hermes-plans/ (all 7 files)
- TRMNL plan already in Done table
- seanime + romm documented in seeds/knowledge.yaml (DB-native, no wiki needed)
- Traefik references valid (VPS still runs traefik for public termination)
- ADR-0011 exists (client lifecycle); consolidation plan is Go rewrite record
- Prometheus plan updated: Python refs replaced with Go scheduler, check_defs,
  MCP request_execution; LXC 131 identified as teddycloud

Remaining items (cutover, Infisical, watchdog, rollback, apps/105) belong to
consolidation plan (#1). 4 of 6 plans now Done.
This commit is contained in:
2026-07-08 11:09:37 +02:00
parent a3ebd12e90
commit 43aaf2a318
4 changed files with 59 additions and 52 deletions

View File

@@ -1,8 +1,8 @@
# Oikos metrics stack — Prometheus LXC (planned)
Lifecycle state: **planned** (see [oikos/ontology.yaml](../oikos/ontology.yaml)
Lifecycle state: **planned** (see [seeds/ontology.yaml](../seeds/ontology.yaml)
lifecycle). No LXC exists yet — this is the plan doc that state requires
before provisioning starts. Do not add a `hosts:` entry with a guessed
before provisioning starts. Do not add a host entry with a guessed
`pve_id` until the LXC is actually created; Proxmox assigns the real ID
at `pct create` time.
@@ -10,19 +10,17 @@ at `pct create` time.
Week-3 reliability layer (see [OIKOS.md](../.agents/OIKOS.md)) wants trend
signals — "disk full in ~9 days at current rate", temperature creep —
which need a real time-series store. The scheduler
([oikos/scheduler.py](../oikos/scheduler.py)) currently does point-in-time
threshold checks only; Prometheus is the one new piece of infrastructure
the 30-day roadmap calls for.
which need a real time-series store. The Go scheduler
([internal/scheduler/scheduler.go](../internal/scheduler/scheduler.go))
currently does point-in-time threshold checks only; Prometheus is the one
new piece of infrastructure the 30-day roadmap calls for.
## Note: undocumented LXC 131 on hubris
## Note: LXC 131 is teddycloud
Oikos's drift detector ([oikos/drift.py](../oikos/drift.py)) found
`pve_id 131` live on hubris (via `pct list`) with no `inventory.yaml`
entry — created outside the provision-node runbook, identity unknown
from this repo. **Investigate what 131 is before assuming any pve_id is
free**; don't let Proxmox auto-assign into a range you haven't confirmed
is actually unused end-to-end.
The drift detector (`drift` check kind in `check_defs`) found
`pve_id 131` live on hubris — this is **teddycloud** (`lxc:teddycloud`),
which is now documented in `seeds/knowledge.yaml`. The pve_id range is
fully accounted for (101-134). Next available VMIDs start at 135.
## Plan
@@ -31,7 +29,7 @@ is actually unused end-to-end.
the Week-4 Oikos Console renders its own sparklines from the Prometheus
HTTP API, per the plan's Week-3 scope decision).
- **Networking:** LAN + mesh-gated only, no public ingress (matches
`homelab_mcp`/`secrets_issuance`'s `MESH_SUBNETS` pattern) — Prometheus
`mcp` / `secrets-issuance` `MESH_SUBNETS` pattern) — Prometheus
exposes host/service metadata that shouldn't be public.
- **Scrape targets:** node_exporter on hubris and strong (Proxmox hosts)
+ any LXC the scheduler needs disk/temp trend data from beyond what
@@ -42,28 +40,32 @@ is actually unused end-to-end.
## Provisioning steps (once pve_id is assigned)
Follow [lifecycle-provision-node](../.agents/skills/lifecycle-provision-node/SKILL.md):
1. `pct create <new-id> ...` on hubris — confirm the assigned ID doesn't
collide with 131 or anything else live.
2. `homelab client add metrics` (or the chosen name) with `state:
provisioning`, stub `containers/<id>-metrics.md`.
1. Add entity to `seeds/inventory.yaml` with state `planned` and the
assigned pve_id. Run `oikos seed` to ingest.
2. `pct create <new-id> ...` on hubris — next available PVE ID is 135.
3. Install Prometheus + node_exporter (Debian package or binary release —
decide at implementation time; no strong preference recorded here).
4. Point node_exporter at hubris + strong (either install locally on each,
or scrape via SSH-tunneled metrics — install locally is simpler and is
the standard approach).
5. Follow [lifecycle-activate-node](../.agents/skills/lifecycle-activate-node/SKILL.md)
to flip to `active`, complete the doc page, regenerate
`hosts/*.yaml` + `infrastructure/topology.md`.
6. Extend `oikos/scheduler.py`'s disk/temp probes to query Prometheus
rate() instead of (or alongside) the current live SSH `df` probe, so
drift/threshold Signals gain trend evidence ("full in ~9 days") instead
of only a point-in-time percentage.
or scrape via SSH-tunneled metrics — install locally is simpler).
5. Transition to `active` via `PATCH /api/v1/entities/{slug}` or
operator approval through the execution flow.
6. Add a `prometheus` check kind to `check_defs` in the scheduler
(`internal/scheduler/scheduler.go`) so trend signals can query
Prometheus `rate()` alongside the current SSH `df` probe, giving
"full in ~9 days" predictions instead of only point-in-time
percentages. Register a `prometheus` check def via `POST /api/v1/checks`.
7. Run `oikos export` to regenerate `seeds/inventory.yaml` for git.
## Open question for the operator
Package choice (apt `prometheus` vs upstream binary release) and exact
scrape interval aren't decided here — pick at implementation time based
on what's easiest to keep patched via the existing `homelab apt-audit`/
`apt-upgrade` fleet tooling if using the Debian package.
on what's easiest to keep patched via MCP `request_execution(action="apt_upgrade")`.
## Changelog
### 2026-07-08 — Go references updated
Replaced Python references (`oikos/scheduler.py`, `oikos/drift.py`,
`bin/homelab`, `homelab client add`) with Go equivalents: `internal/scheduler/`,
`check_defs` drift kind, MCP `request_execution`, `seeds/inventory.yaml`.
LXC 131 identified as teddycloud (PVE range 101-134 fully accounted for).