Files
oikos/archive/hermes-plans/2026-06-03_150000-homelab-structure-revision.md
dtoro 6b75f7302d db as source of truth: wiki→seeds, archive old artifacts, knowledge ingestion
- Migrations 010 (content_hash) + 011 (search tsvector column)
- new: internal/knowledge/seed.go — knowledge seed ingest engine
- new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge
- wire knowledge ingest into oikos seed pipeline
- convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml
- archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/
- delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py
- remove empty knowledge/ directory tree
2026-07-07 20:22:30 +02:00

12 KiB

Homelab structure revision & improvement plan

Goal

Identify structural issues in the current hubris homelab topology and propose an actionable improvement roadmap — DNS consolidation, monitoring gaps, backup recovery, mesh completion, resource rightsizing, and operational hygiene.


Current state summary

Dimension Status
Hypervisor hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24)
LXCs 14 active, 1 retired (124), 1 new (107 dns)
VMs HAOS (108), ZimaOS (100)
Workstations mac-mini (macOS), republic-laptop, ludo-mini
VPS 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn
Switch SODOLA 5-Port 2.5Gbit (L2, flat)
Networking 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing
Mesh Netbird + Tailscale (migrating), ~mixed state
DNS 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS
Backups DISABLED since 2026-04-22
Monitoring claudio-monitor (host health only, no apt/docker checks yet)
Agent enrollment 3/19 hosts enrolled (hubris, apps, republic-laptop)

Issues identified

1. Three overlapping DNS sources (highest risk)

Problem: Technitium on CT 107, NetBird managed DNS zone, and the public IONOS wildcard all answer *.hubris.network queries. The dns.md doc still references the old dnsmasq on LXC 124 (though the change log says it moved). NetBird's managed DNS bypasses Technitium entirely for some app names — there is no single source of truth for DNS.

Risk: Mismatched answers → services unreachable → "works on some clients but not others" debugging sessions. Already cost time when auth.hubris.network re-pointed to the VPS.

Proposal:

  • Phase out NetBird managed DNS zone for hubris.network — Technitium is the authoritative answerer for mesh & LAN clients
  • Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale MagicDNS fallbacks)
  • Document the full authoritative chain: Technitium → upstream forwarders → public
  • Track Technitium config in git (dtoro/technitium-config or equivalent)

2. Backups disabled with no alternative (data loss risk)

Problem: The only backup was restic to an external USB that caused host crashes. It was disabled 2026-04-22 as an A/B test — host stability was confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The drive is still removed.

Risk: /mnt/library (~429 GB of irreplaceable data: photos, documents, git repos, gitea data) has no off-host copy. Single-disk failure = total loss.

Proposal:

  • Re-evaluate the USB drive stability with the new SODOLA switch topology (direct rear USB 3.0 port, no hub chain)
  • OR adopt a cloud-backed strategy: rclone-to-Hetzner Storage Box or Backblaze B2 for the irreplaceable subset (docs, photos, gitea data)
  • OR use hubris's own zfs send to a second host/disk if ZFS is feasible
  • Minimum viable: at minimum restore gitea backups + sops-encrypted secrets via an off-site cron (cheap B2 bucket)

3. Mesh migration still incomplete

Problem: Most LXCs still use Tailscale. This forces per-LXC DNS workarounds (/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel. Mesh migration doc (mesh.md) is comprehensive but execution stalled.

Proposal:

  • Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
  • Remove Tailscale from the PVE host
  • Ensure all LXCs resolve *.hubris.network via Technitium (no more /etc/hosts overrides)
  • Document Netbird client on each LXC (netbird version, setup key rotation)

4. LXC resource imbalance & disk pressure

Problem:

LXC Cores RAM Rootfs Disk usage
mule-images (120) 6 12 GiB 60 GiB photo AI, reasonable
arriman (122) 4 8 GiB 24 GiB media stack, OK
nextcloud (114) 4 6 GiB 25 GiB file sync, OK
apps (105) 2 4 GiB 30 GiB 6+ services, tight
paperless (103) 2 3 GiB 8 GiB 86.9% disk
elementsynapse (118) 1 2 GiB 8 GiB 86.8% disk
gitea (104) 1 1 GiB 8 GiB git server, adequate
caddy (121) 1 512 MiB 6 GiB fine for reverse proxy

Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105) is undersized for 6+ services.

Proposal:

  • Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
  • Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
  • Enable claudio-monitor's disk check to alert before next crisis

5. VPS is a single point of failure

Problem: One IONOS VM runs netbird management (control plane), traefik (public ingress), authentik (identity), and coturn (TURN relay). If it goes down: no remote mesh, no public services, no auth.

Proposal:

  • Document a VPS recovery runbook (how to restore from a known-working backup)
  • Consider splitting authentik into a separate host or at minimum having a standby configuration
  • Not a high priority (the VPS has been stable) but worth documenting the blast radius and recovery path

6. No centralized logging

Problem: Each LXC has independent journald. Cross-service debugging involves hopping between pct exec <id> -- journalctl -u <service>. There is no aggregation or retention.

Proposal:

  • Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
  • Ship logs to a central Loki instance on apps (105) or a new small LXC
  • Grafana dashboard optional — even a simple logcli query saves time

7. Agent enrollment incomplete

Problem: Only hubris, apps, and republic-laptop are enrolled in the homelab-context system (age keys, sync timers, MCP access). mac-mini, ludo-mini, claudio-bot, and all other LXCs are not.

Proposal:

  • Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
  • Enroll mac-mini (macOS — exercises the launchd timer path)
  • Enroll ludo-mini (needs SSH user config in inventory first)
  • Wire claudio-bot into inventory-aware queries

8. Configuration drift on untracked configs

Problem: Technitium config, dnsmasq (legacy), and several service-specific configs are not git-tracked.

Proposal:

  • Track Technitium zone backup + compose config in a git repo
  • Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy

9. No capacity planning / resource monitoring

Problem: No trend data on CPU, RAM, or disk growth. The 86% disk alerts were discovered reactively. Rootfs resize is painful (requires Proxmox stop + resize + growfs inside).

Proposal:

  • Enable the missing claudio-monitor checks (disk growth trend, apt upgradable counts, docker image drift)
  • Set up a simple Prometheus + node_exporter on hubris or use the PVE API directly
  • At minimum, surface disk usage in the existing homelab-mcp management tools

10. No standard deploy / orchestration for bare-metal LXCs

Problem: Some services are bare-metal CLI apps (sophia, claudio-bot), some are Docker on apps (105), some are Portainer-managed. No consistent deploy pattern means every new service reinvents the deployment.

Proposal:

  • Don't over-engineer this — the current pragmatism works
  • Just document the decision tree:
    • Needs /mnt/library mount + heavy I/O → dedicated LXC
    • Small stateless web service → Docker on apps (105)
    • Media stack → dedicated LXC (arriman, jellyfin)
    • Everything else → judge by complexity

Phased implementation plan

Phase 1 — Critical fixes (this week)

  1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
  2. Enable claudio-monitor disk check + disk-growth alerting
  3. Pick one backup strategy and implement minimum viable (e.g. nightly gitea dump + sops-encrypted secrets to B2 via rclone)
  4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone for hubris.network)

Phase 2 — Mesh consolidation (next week)

  1. Batch-migrate remaining LXCs from Tailscale to Netbird
  2. Remove Tailscale from PVE host
  3. Remove all per-LXC /etc/hosts DNS overrides
  4. Update DNS documentation to reflect Technitium as single source

Phase 3 — Agent enrollment & logging (next 2 weeks)

  1. Enroll all LXCs in homelab-context (age keys, sync timers)
  2. Enroll mac-mini (macOS launchd path — exercises untested code path)
  3. Enroll ludo-mini
  4. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs

Phase 4 — Resource & monitoring hardening (next month)

  1. Resize apps (105) rootfs, bump RAM
  2. Deploy Prometheus + node_exporter or equivalent for trend data
  3. Track Technitium config in git with auto-deploy
  4. Write VPS recovery runbook

Phase 5 — Drive re-evaluation (optional, behind host-stability gate)

  1. Re-attach USB backup drive with the new SODOLA topology (direct port)
  2. If stable for 7 days, re-enable restic backup schedule (chunked)
  3. If not stable, finalize cloud backup as permanent strategy

Files likely to change

Path Change
/opt/homelab-context/inventory.yaml LXCs enrolled, resource updates
/opt/homelab-context/secrets/*.yaml New recipients for enrolled LXCs
/opt/homelab-context/infrastructure/dns.md Reflect Technitium as sole source
/opt/homelab-context/infrastructure/mesh.md Remove Tailscale references post-migration
/opt/homelab-context/infrastructure/backups.md New strategy
/opt/homelab-context/infrastructure/monitoring.md Enable missing checks
/opt/homelab-context/containers/103-paperless.md Rootfs resize
/opt/homelab-context/containers/118-elementsynapse.md Rootfs resize
/opt/homelab-context/containers/105-apps.md Resource bump, logging addition
/opt/homelab-context/containers/index.md Updated resource table
.sops.yaml New age pubkeys for enrolled LXCs

Verification

Each phase ends with a verification milestone:

  • Phase 1: claudio-monitor triggers on paperless disk → confirmed alert. Backup of gitea data lands in B2 (or equivalent). DNS query from any LXC returns Technitium answer.
  • Phase 2: netbird status shows all LXCs connected. Tailscale not running on PVE host. curl auth.hubris.network from any LXC resolves correctly without /etc/hosts.
  • Phase 3: Every LXC has /opt/homelab-context/ + /etc/age/key.txt. MCP tools return valid host info for all enrolled LXC names. journalctl shows promtail shipping to Loki.
  • Phase 4: claudio-monitor shows disk growth trend. apps (105) can run all 6+ services without OOM.

Risks & tradeoffs

  • Netbird migration window: All LXCs will briefly lose mesh connectivity during the Tailscale→Netbird cutover. Schedule in off-hours.
  • Backup cost: B2/e2 costs ~$5/month for ~500 GB. The USB drive was free but unstable — trade money for reliability.
  • DNS consolidation: Removing the NetBird managed DNS zone means any NetBird-specific names stop resolving for hubris.network — verify nothing depends on that path.
  • Loki on apps (105): Adds another container to an already-loaded host. May need to bump resources before deploying.
  • Agent enrollment on every LXC: Each enrollment creates an age keypair and commits a pubkey to inventory. Process is scriptable via homelab client add but still takes ~2 min per host for verification.

Open questions

  1. Is the USB backup drive still physically attached to hubris? If not, the simplest "re-enable" path requires physically re-attaching it.
  2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or was it fully decommissioned? The dns.md changelog says "shut down" but index.md lists it as "running".
  3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure — needed to plan rootfs resizes safely.
  4. Does the user want to keep Tailscale on any host for a specific reason, or is full Netbird migration the clear goal?