Files
oikos/archive/hermes-plans/2026-06-03_150000-homelab-structure-revision.md
dtoro 6b75f7302d db as source of truth: wiki→seeds, archive old artifacts, knowledge ingestion
- Migrations 010 (content_hash) + 011 (search tsvector column)
- new: internal/knowledge/seed.go — knowledge seed ingest engine
- new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge
- wire knowledge ingest into oikos seed pipeline
- convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml
- archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/
- delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py
- remove empty knowledge/ directory tree
2026-07-07 20:22:30 +02:00

271 lines
12 KiB
Markdown

# Homelab structure revision & improvement plan
## Goal
Identify structural issues in the current hubris homelab topology and propose an
actionable improvement roadmap — DNS consolidation, monitoring gaps, backup
recovery, mesh completion, resource rightsizing, and operational hygiene.
---
## Current state summary
| Dimension | Status |
|-----------|--------|
| Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) |
| LXCs | 14 active, 1 retired (124), 1 new (107 dns) |
| VMs | HAOS (108), ZimaOS (100) |
| Workstations | mac-mini (macOS), republic-laptop, ludo-mini |
| VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn |
| Switch | SODOLA 5-Port 2.5Gbit (L2, flat) |
| Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing |
| Mesh | Netbird + Tailscale (migrating), ~mixed state |
| DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS |
| Backups | DISABLED since 2026-04-22 |
| Monitoring | claudio-monitor (host health only, no apt/docker checks yet) |
| Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) |
---
## Issues identified
### 1. Three overlapping DNS sources (highest risk)
**Problem:** Technitium on CT 107, NetBird managed DNS zone, and the public
IONOS wildcard all answer `*.hubris.network` queries. The dns.md doc still
references the old dnsmasq on LXC 124 (though the change log says it moved).
NetBird's managed DNS bypasses Technitium entirely for some app names — there
is no single source of truth for DNS.
**Risk:** Mismatched answers → services unreachable → "works on some clients
but not others" debugging sessions. Already cost time when `auth.hubris.network`
re-pointed to the VPS.
**Proposal:**
- Phase out NetBird managed DNS zone for `hubris.network` — Technitium is the
authoritative answerer for mesh & LAN clients
- Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale
MagicDNS fallbacks)
- Document the full authoritative chain: Technitium → upstream forwarders → public
- Track Technitium config in git (dtoro/technitium-config or equivalent)
### 2. Backups disabled with no alternative (data loss risk)
**Problem:** The only backup was restic to an external USB that caused host
crashes. It was disabled 2026-04-22 as an A/B test — host stability was
confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The
drive is still removed.
**Risk:** `/mnt/library` (~429 GB of irreplaceable data: photos, documents,
git repos, gitea data) has no off-host copy. Single-disk failure = total loss.
**Proposal:**
- Re-evaluate the USB drive stability with the new SODOLA switch topology
(direct rear USB 3.0 port, no hub chain)
- OR adopt a cloud-backed strategy: `rclone`-to-Hetzner Storage Box or
Backblaze B2 for the irreplaceable subset (docs, photos, gitea data)
- OR use hubris's own `zfs send` to a second host/disk if ZFS is feasible
- Minimum viable: at minimum restore gitea backups + sops-encrypted secrets
via an off-site cron (cheap B2 bucket)
### 3. Mesh migration still incomplete
**Problem:** Most LXCs still use Tailscale. This forces per-LXC DNS workarounds
(/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel.
Mesh migration doc (mesh.md) is comprehensive but execution stalled.
**Proposal:**
- Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
- Remove Tailscale from the PVE host
- Ensure all LXCs resolve `*.hubris.network` via Technitium (no more /etc/hosts
overrides)
- Document Netbird client on each LXC (netbird version, setup key rotation)
### 4. LXC resource imbalance & disk pressure
**Problem:**
| LXC | Cores | RAM | Rootfs | Disk usage |
|-----|-------|-----|--------|------------|
| mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable |
| arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK |
| nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK |
| apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight |
| paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk |
| elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk |
| gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate |
| caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy |
Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105)
is undersized for 6+ services.
**Proposal:**
- Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
- Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
- Enable claudio-monitor's disk check to alert before next crisis
### 5. VPS is a single point of failure
**Problem:** One IONOS VM runs netbird management (control plane), traefik
(public ingress), authentik (identity), and coturn (TURN relay). If it goes
down: no remote mesh, no public services, no auth.
**Proposal:**
- Document a VPS recovery runbook (how to restore from a known-working backup)
- Consider splitting authentik into a separate host or at minimum having a
standby configuration
- Not a high priority (the VPS has been stable) but worth documenting the
blast radius and recovery path
### 6. No centralized logging
**Problem:** Each LXC has independent journald. Cross-service debugging
involves hopping between `pct exec <id> -- journalctl -u <service>`. There is
no aggregation or retention.
**Proposal:**
- Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
- Ship logs to a central Loki instance on apps (105) or a new small LXC
- Grafana dashboard optional — even a simple `logcli` query saves time
### 7. Agent enrollment incomplete
**Problem:** Only hubris, apps, and republic-laptop are enrolled in the
homelab-context system (age keys, sync timers, MCP access). mac-mini,
ludo-mini, claudio-bot, and all other LXCs are not.
**Proposal:**
- Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
- Enroll mac-mini (macOS — exercises the launchd timer path)
- Enroll ludo-mini (needs SSH user config in inventory first)
- Wire claudio-bot into inventory-aware queries
### 8. Configuration drift on untracked configs
**Problem:** Technitium config, dnsmasq (legacy), and several service-specific
configs are not git-tracked.
**Proposal:**
- Track Technitium zone backup + compose config in a git repo
- Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy
### 9. No capacity planning / resource monitoring
**Problem:** No trend data on CPU, RAM, or disk growth. The 86% disk alerts
were discovered reactively. Rootfs resize is painful (requires Proxmox stop +
resize + growfs inside).
**Proposal:**
- Enable the missing claudio-monitor checks (disk growth trend, apt upgradable
counts, docker image drift)
- Set up a simple Prometheus + node_exporter on hubris or use the PVE API
directly
- At minimum, surface disk usage in the existing homelab-mcp management tools
### 10. No standard deploy / orchestration for bare-metal LXCs
**Problem:** Some services are bare-metal CLI apps (sophia, claudio-bot),
some are Docker on apps (105), some are Portainer-managed. No consistent
deploy pattern means every new service reinvents the deployment.
**Proposal:**
- Don't over-engineer this — the current pragmatism works
- Just document the decision tree:
- Needs `/mnt/library` mount + heavy I/O → dedicated LXC
- Small stateless web service → Docker on apps (105)
- Media stack → dedicated LXC (arriman, jellyfin)
- Everything else → judge by complexity
---
## Phased implementation plan
### Phase 1 — Critical fixes (this week)
1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
2. Enable claudio-monitor disk check + disk-growth alerting
3. Pick one backup strategy and implement minimum viable (e.g. nightly
gitea dump + sops-encrypted secrets to B2 via rclone)
4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone
for hubris.network)
### Phase 2 — Mesh consolidation (next week)
5. Batch-migrate remaining LXCs from Tailscale to Netbird
6. Remove Tailscale from PVE host
7. Remove all per-LXC /etc/hosts DNS overrides
8. Update DNS documentation to reflect Technitium as single source
### Phase 3 — Agent enrollment & logging (next 2 weeks)
9. Enroll all LXCs in homelab-context (age keys, sync timers)
10. Enroll mac-mini (macOS launchd path — exercises untested code path)
11. Enroll ludo-mini
12. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs
### Phase 4 — Resource & monitoring hardening (next month)
13. Resize apps (105) rootfs, bump RAM
14. Deploy Prometheus + node_exporter or equivalent for trend data
15. Track Technitium config in git with auto-deploy
16. Write VPS recovery runbook
### Phase 5 — Drive re-evaluation (optional, behind host-stability gate)
17. Re-attach USB backup drive with the new SODOLA topology (direct port)
18. If stable for 7 days, re-enable restic backup schedule (chunked)
19. If not stable, finalize cloud backup as permanent strategy
---
## Files likely to change
| Path | Change |
|------|--------|
| `/opt/homelab-context/inventory.yaml` | LXCs enrolled, resource updates |
| `/opt/homelab-context/secrets/*.yaml` | New recipients for enrolled LXCs |
| `/opt/homelab-context/infrastructure/dns.md` | Reflect Technitium as sole source |
| `/opt/homelab-context/infrastructure/mesh.md` | Remove Tailscale references post-migration |
| `/opt/homelab-context/infrastructure/backups.md` | New strategy |
| `/opt/homelab-context/infrastructure/monitoring.md` | Enable missing checks |
| `/opt/homelab-context/containers/103-paperless.md` | Rootfs resize |
| `/opt/homelab-context/containers/118-elementsynapse.md` | Rootfs resize |
| `/opt/homelab-context/containers/105-apps.md` | Resource bump, logging addition |
| `/opt/homelab-context/containers/index.md` | Updated resource table |
| `.sops.yaml` | New age pubkeys for enrolled LXCs |
## Verification
Each phase ends with a verification milestone:
- Phase 1: `claudio-monitor` triggers on paperless disk → confirmed alert. Backup
of gitea data lands in B2 (or equivalent). DNS query from any LXC returns
Technitium answer.
- Phase 2: `netbird status` shows all LXCs connected. Tailscale not running on
PVE host. `curl auth.hubris.network` from any LXC resolves correctly without
/etc/hosts.
- Phase 3: Every LXC has `/opt/homelab-context/` + `/etc/age/key.txt`. MCP
tools return valid host info for all enrolled LXC names. `journalctl` shows
promtail shipping to Loki.
- Phase 4: `claudio-monitor` shows disk growth trend. apps (105) can run all
6+ services without OOM.
## Risks & tradeoffs
- **Netbird migration window:** All LXCs will briefly lose mesh connectivity
during the Tailscale→Netbird cutover. Schedule in off-hours.
- **Backup cost:** B2/e2 costs ~$5/month for ~500 GB. The USB drive was free
but unstable — trade money for reliability.
- **DNS consolidation:** Removing the NetBird managed DNS zone means any
NetBird-specific names stop resolving for hubris.network — verify nothing
depends on that path.
- **Loki on apps (105):** Adds another container to an already-loaded host.
May need to bump resources before deploying.
- **Agent enrollment on every LXC:** Each enrollment creates an age keypair
and commits a pubkey to inventory. Process is scriptable via `homelab client
add` but still takes ~2 min per host for verification.
## Open questions
1. Is the USB backup drive still physically attached to hubris? If not, the
simplest "re-enable" path requires physically re-attaching it.
2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or
was it fully decommissioned? The dns.md changelog says "shut down" but
index.md lists it as "running".
3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure
— needed to plan rootfs resizes safely.
4. Does the user want to keep Tailscale on any host for a specific reason, or
is full Netbird migration the clear goal?