# Homelab structure revision & improvement plan ## Goal Identify structural issues in the current hubris homelab topology and propose an actionable improvement roadmap — DNS consolidation, monitoring gaps, backup recovery, mesh completion, resource rightsizing, and operational hygiene. --- ## Current state summary | Dimension | Status | |-----------|--------| | Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) | | LXCs | 14 active, 1 retired (124), 1 new (107 dns) | | VMs | HAOS (108), ZimaOS (100) | | Workstations | mac-mini (macOS), republic-laptop, ludo-mini | | VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn | | Switch | SODOLA 5-Port 2.5Gbit (L2, flat) | | Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing | | Mesh | Netbird + Tailscale (migrating), ~mixed state | | DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS | | Backups | DISABLED since 2026-04-22 | | Monitoring | claudio-monitor (host health only, no apt/docker checks yet) | | Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) | --- ## Issues identified ### 1. Three overlapping DNS sources (highest risk) **Problem:** Technitium on CT 107, NetBird managed DNS zone, and the public IONOS wildcard all answer `*.hubris.network` queries. The dns.md doc still references the old dnsmasq on LXC 124 (though the change log says it moved). NetBird's managed DNS bypasses Technitium entirely for some app names — there is no single source of truth for DNS. **Risk:** Mismatched answers → services unreachable → "works on some clients but not others" debugging sessions. Already cost time when `auth.hubris.network` re-pointed to the VPS. **Proposal:** - Phase out NetBird managed DNS zone for `hubris.network` — Technitium is the authoritative answerer for mesh & LAN clients - Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale MagicDNS fallbacks) - Document the full authoritative chain: Technitium → upstream forwarders → public - Track Technitium config in git (dtoro/technitium-config or equivalent) ### 2. Backups disabled with no alternative (data loss risk) **Problem:** The only backup was restic to an external USB that caused host crashes. It was disabled 2026-04-22 as an A/B test — host stability was confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The drive is still removed. **Risk:** `/mnt/library` (~429 GB of irreplaceable data: photos, documents, git repos, gitea data) has no off-host copy. Single-disk failure = total loss. **Proposal:** - Re-evaluate the USB drive stability with the new SODOLA switch topology (direct rear USB 3.0 port, no hub chain) - OR adopt a cloud-backed strategy: `rclone`-to-Hetzner Storage Box or Backblaze B2 for the irreplaceable subset (docs, photos, gitea data) - OR use hubris's own `zfs send` to a second host/disk if ZFS is feasible - Minimum viable: at minimum restore gitea backups + sops-encrypted secrets via an off-site cron (cheap B2 bucket) ### 3. Mesh migration still incomplete **Problem:** Most LXCs still use Tailscale. This forces per-LXC DNS workarounds (/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel. Mesh migration doc (mesh.md) is comprehensive but execution stalled. **Proposal:** - Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window - Remove Tailscale from the PVE host - Ensure all LXCs resolve `*.hubris.network` via Technitium (no more /etc/hosts overrides) - Document Netbird client on each LXC (netbird version, setup key rotation) ### 4. LXC resource imbalance & disk pressure **Problem:** | LXC | Cores | RAM | Rootfs | Disk usage | |-----|-------|-----|--------|------------| | mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable | | arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK | | nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK | | apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight | | paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk | | elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk | | gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate | | caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy | Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105) is undersized for 6+ services. **Proposal:** - Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB) - Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs - Enable claudio-monitor's disk check to alert before next crisis ### 5. VPS is a single point of failure **Problem:** One IONOS VM runs netbird management (control plane), traefik (public ingress), authentik (identity), and coturn (TURN relay). If it goes down: no remote mesh, no public services, no auth. **Proposal:** - Document a VPS recovery runbook (how to restore from a known-working backup) - Consider splitting authentik into a separate host or at minimum having a standby configuration - Not a high priority (the VPS has been stable) but worth documenting the blast radius and recovery path ### 6. No centralized logging **Problem:** Each LXC has independent journald. Cross-service debugging involves hopping between `pct exec -- journalctl -u `. There is no aggregation or retention. **Proposal:** - Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC - Ship logs to a central Loki instance on apps (105) or a new small LXC - Grafana dashboard optional — even a simple `logcli` query saves time ### 7. Agent enrollment incomplete **Problem:** Only hubris, apps, and republic-laptop are enrolled in the homelab-context system (age keys, sync timers, MCP access). mac-mini, ludo-mini, claudio-bot, and all other LXCs are not. **Proposal:** - Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.) - Enroll mac-mini (macOS — exercises the launchd timer path) - Enroll ludo-mini (needs SSH user config in inventory first) - Wire claudio-bot into inventory-aware queries ### 8. Configuration drift on untracked configs **Problem:** Technitium config, dnsmasq (legacy), and several service-specific configs are not git-tracked. **Proposal:** - Track Technitium zone backup + compose config in a git repo - Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy ### 9. No capacity planning / resource monitoring **Problem:** No trend data on CPU, RAM, or disk growth. The 86% disk alerts were discovered reactively. Rootfs resize is painful (requires Proxmox stop + resize + growfs inside). **Proposal:** - Enable the missing claudio-monitor checks (disk growth trend, apt upgradable counts, docker image drift) - Set up a simple Prometheus + node_exporter on hubris or use the PVE API directly - At minimum, surface disk usage in the existing homelab-mcp management tools ### 10. No standard deploy / orchestration for bare-metal LXCs **Problem:** Some services are bare-metal CLI apps (sophia, claudio-bot), some are Docker on apps (105), some are Portainer-managed. No consistent deploy pattern means every new service reinvents the deployment. **Proposal:** - Don't over-engineer this — the current pragmatism works - Just document the decision tree: - Needs `/mnt/library` mount + heavy I/O → dedicated LXC - Small stateless web service → Docker on apps (105) - Media stack → dedicated LXC (arriman, jellyfin) - Everything else → judge by complexity --- ## Phased implementation plan ### Phase 1 — Critical fixes (this week) 1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB 2. Enable claudio-monitor disk check + disk-growth alerting 3. Pick one backup strategy and implement minimum viable (e.g. nightly gitea dump + sops-encrypted secrets to B2 via rclone) 4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone for hubris.network) ### Phase 2 — Mesh consolidation (next week) 5. Batch-migrate remaining LXCs from Tailscale to Netbird 6. Remove Tailscale from PVE host 7. Remove all per-LXC /etc/hosts DNS overrides 8. Update DNS documentation to reflect Technitium as single source ### Phase 3 — Agent enrollment & logging (next 2 weeks) 9. Enroll all LXCs in homelab-context (age keys, sync timers) 10. Enroll mac-mini (macOS launchd path — exercises untested code path) 11. Enroll ludo-mini 12. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs ### Phase 4 — Resource & monitoring hardening (next month) 13. Resize apps (105) rootfs, bump RAM 14. Deploy Prometheus + node_exporter or equivalent for trend data 15. Track Technitium config in git with auto-deploy 16. Write VPS recovery runbook ### Phase 5 — Drive re-evaluation (optional, behind host-stability gate) 17. Re-attach USB backup drive with the new SODOLA topology (direct port) 18. If stable for 7 days, re-enable restic backup schedule (chunked) 19. If not stable, finalize cloud backup as permanent strategy --- ## Files likely to change | Path | Change | |------|--------| | `/opt/homelab-context/inventory.yaml` | LXCs enrolled, resource updates | | `/opt/homelab-context/secrets/*.yaml` | New recipients for enrolled LXCs | | `/opt/homelab-context/infrastructure/dns.md` | Reflect Technitium as sole source | | `/opt/homelab-context/infrastructure/mesh.md` | Remove Tailscale references post-migration | | `/opt/homelab-context/infrastructure/backups.md` | New strategy | | `/opt/homelab-context/infrastructure/monitoring.md` | Enable missing checks | | `/opt/homelab-context/containers/103-paperless.md` | Rootfs resize | | `/opt/homelab-context/containers/118-elementsynapse.md` | Rootfs resize | | `/opt/homelab-context/containers/105-apps.md` | Resource bump, logging addition | | `/opt/homelab-context/containers/index.md` | Updated resource table | | `.sops.yaml` | New age pubkeys for enrolled LXCs | ## Verification Each phase ends with a verification milestone: - Phase 1: `claudio-monitor` triggers on paperless disk → confirmed alert. Backup of gitea data lands in B2 (or equivalent). DNS query from any LXC returns Technitium answer. - Phase 2: `netbird status` shows all LXCs connected. Tailscale not running on PVE host. `curl auth.hubris.network` from any LXC resolves correctly without /etc/hosts. - Phase 3: Every LXC has `/opt/homelab-context/` + `/etc/age/key.txt`. MCP tools return valid host info for all enrolled LXC names. `journalctl` shows promtail shipping to Loki. - Phase 4: `claudio-monitor` shows disk growth trend. apps (105) can run all 6+ services without OOM. ## Risks & tradeoffs - **Netbird migration window:** All LXCs will briefly lose mesh connectivity during the Tailscale→Netbird cutover. Schedule in off-hours. - **Backup cost:** B2/e2 costs ~$5/month for ~500 GB. The USB drive was free but unstable — trade money for reliability. - **DNS consolidation:** Removing the NetBird managed DNS zone means any NetBird-specific names stop resolving for hubris.network — verify nothing depends on that path. - **Loki on apps (105):** Adds another container to an already-loaded host. May need to bump resources before deploying. - **Agent enrollment on every LXC:** Each enrollment creates an age keypair and commits a pubkey to inventory. Process is scriptable via `homelab client add` but still takes ~2 min per host for verification. ## Open questions 1. Is the USB backup drive still physically attached to hubris? If not, the simplest "re-enable" path requires physically re-attaching it. 2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or was it fully decommissioned? The dns.md changelog says "shut down" but index.md lists it as "running". 3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure — needed to plan rootfs resizes safely. 4. Does the user want to keep Tailscale on any host for a specific reason, or is full Netbird migration the clear goal?