- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
12 KiB
Homelab structure revision & improvement plan
Goal
Identify structural issues in the current hubris homelab topology and propose an actionable improvement roadmap — DNS consolidation, monitoring gaps, backup recovery, mesh completion, resource rightsizing, and operational hygiene.
Current state summary
| Dimension | Status |
|---|---|
| Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) |
| LXCs | 14 active, 1 retired (124), 1 new (107 dns) |
| VMs | HAOS (108), ZimaOS (100) |
| Workstations | mac-mini (macOS), republic-laptop, ludo-mini |
| VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn |
| Switch | SODOLA 5-Port 2.5Gbit (L2, flat) |
| Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing |
| Mesh | Netbird + Tailscale (migrating), ~mixed state |
| DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS |
| Backups | DISABLED since 2026-04-22 |
| Monitoring | claudio-monitor (host health only, no apt/docker checks yet) |
| Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) |
Issues identified
1. Three overlapping DNS sources (highest risk)
Problem: Technitium on CT 107, NetBird managed DNS zone, and the public
IONOS wildcard all answer *.hubris.network queries. The dns.md doc still
references the old dnsmasq on LXC 124 (though the change log says it moved).
NetBird's managed DNS bypasses Technitium entirely for some app names — there
is no single source of truth for DNS.
Risk: Mismatched answers → services unreachable → "works on some clients
but not others" debugging sessions. Already cost time when auth.hubris.network
re-pointed to the VPS.
Proposal:
- Phase out NetBird managed DNS zone for
hubris.network— Technitium is the authoritative answerer for mesh & LAN clients - Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale MagicDNS fallbacks)
- Document the full authoritative chain: Technitium → upstream forwarders → public
- Track Technitium config in git (dtoro/technitium-config or equivalent)
2. Backups disabled with no alternative (data loss risk)
Problem: The only backup was restic to an external USB that caused host crashes. It was disabled 2026-04-22 as an A/B test — host stability was confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The drive is still removed.
Risk: /mnt/library (~429 GB of irreplaceable data: photos, documents,
git repos, gitea data) has no off-host copy. Single-disk failure = total loss.
Proposal:
- Re-evaluate the USB drive stability with the new SODOLA switch topology (direct rear USB 3.0 port, no hub chain)
- OR adopt a cloud-backed strategy:
rclone-to-Hetzner Storage Box or Backblaze B2 for the irreplaceable subset (docs, photos, gitea data) - OR use hubris's own
zfs sendto a second host/disk if ZFS is feasible - Minimum viable: at minimum restore gitea backups + sops-encrypted secrets via an off-site cron (cheap B2 bucket)
3. Mesh migration still incomplete
Problem: Most LXCs still use Tailscale. This forces per-LXC DNS workarounds (/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel. Mesh migration doc (mesh.md) is comprehensive but execution stalled.
Proposal:
- Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
- Remove Tailscale from the PVE host
- Ensure all LXCs resolve
*.hubris.networkvia Technitium (no more /etc/hosts overrides) - Document Netbird client on each LXC (netbird version, setup key rotation)
4. LXC resource imbalance & disk pressure
Problem:
| LXC | Cores | RAM | Rootfs | Disk usage |
|---|---|---|---|---|
| mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable |
| arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK |
| nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK |
| apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight |
| paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk |
| elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk |
| gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate |
| caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy |
Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105) is undersized for 6+ services.
Proposal:
- Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
- Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
- Enable claudio-monitor's disk check to alert before next crisis
5. VPS is a single point of failure
Problem: One IONOS VM runs netbird management (control plane), traefik (public ingress), authentik (identity), and coturn (TURN relay). If it goes down: no remote mesh, no public services, no auth.
Proposal:
- Document a VPS recovery runbook (how to restore from a known-working backup)
- Consider splitting authentik into a separate host or at minimum having a standby configuration
- Not a high priority (the VPS has been stable) but worth documenting the blast radius and recovery path
6. No centralized logging
Problem: Each LXC has independent journald. Cross-service debugging
involves hopping between pct exec <id> -- journalctl -u <service>. There is
no aggregation or retention.
Proposal:
- Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
- Ship logs to a central Loki instance on apps (105) or a new small LXC
- Grafana dashboard optional — even a simple
logcliquery saves time
7. Agent enrollment incomplete
Problem: Only hubris, apps, and republic-laptop are enrolled in the homelab-context system (age keys, sync timers, MCP access). mac-mini, ludo-mini, claudio-bot, and all other LXCs are not.
Proposal:
- Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
- Enroll mac-mini (macOS — exercises the launchd timer path)
- Enroll ludo-mini (needs SSH user config in inventory first)
- Wire claudio-bot into inventory-aware queries
8. Configuration drift on untracked configs
Problem: Technitium config, dnsmasq (legacy), and several service-specific configs are not git-tracked.
Proposal:
- Track Technitium zone backup + compose config in a git repo
- Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy
9. No capacity planning / resource monitoring
Problem: No trend data on CPU, RAM, or disk growth. The 86% disk alerts were discovered reactively. Rootfs resize is painful (requires Proxmox stop + resize + growfs inside).
Proposal:
- Enable the missing claudio-monitor checks (disk growth trend, apt upgradable counts, docker image drift)
- Set up a simple Prometheus + node_exporter on hubris or use the PVE API directly
- At minimum, surface disk usage in the existing homelab-mcp management tools
10. No standard deploy / orchestration for bare-metal LXCs
Problem: Some services are bare-metal CLI apps (sophia, claudio-bot), some are Docker on apps (105), some are Portainer-managed. No consistent deploy pattern means every new service reinvents the deployment.
Proposal:
- Don't over-engineer this — the current pragmatism works
- Just document the decision tree:
- Needs
/mnt/librarymount + heavy I/O → dedicated LXC - Small stateless web service → Docker on apps (105)
- Media stack → dedicated LXC (arriman, jellyfin)
- Everything else → judge by complexity
- Needs
Phased implementation plan
Phase 1 — Critical fixes (this week)
- Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
- Enable claudio-monitor disk check + disk-growth alerting
- Pick one backup strategy and implement minimum viable (e.g. nightly gitea dump + sops-encrypted secrets to B2 via rclone)
- Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone for hubris.network)
Phase 2 — Mesh consolidation (next week)
- Batch-migrate remaining LXCs from Tailscale to Netbird
- Remove Tailscale from PVE host
- Remove all per-LXC /etc/hosts DNS overrides
- Update DNS documentation to reflect Technitium as single source
Phase 3 — Agent enrollment & logging (next 2 weeks)
- Enroll all LXCs in homelab-context (age keys, sync timers)
- Enroll mac-mini (macOS launchd path — exercises untested code path)
- Enroll ludo-mini
- Deploy log shipper (Loki + promtail) on apps (105) + all LXCs
Phase 4 — Resource & monitoring hardening (next month)
- Resize apps (105) rootfs, bump RAM
- Deploy Prometheus + node_exporter or equivalent for trend data
- Track Technitium config in git with auto-deploy
- Write VPS recovery runbook
Phase 5 — Drive re-evaluation (optional, behind host-stability gate)
- Re-attach USB backup drive with the new SODOLA topology (direct port)
- If stable for 7 days, re-enable restic backup schedule (chunked)
- If not stable, finalize cloud backup as permanent strategy
Files likely to change
| Path | Change |
|---|---|
/opt/homelab-context/inventory.yaml |
LXCs enrolled, resource updates |
/opt/homelab-context/secrets/*.yaml |
New recipients for enrolled LXCs |
/opt/homelab-context/infrastructure/dns.md |
Reflect Technitium as sole source |
/opt/homelab-context/infrastructure/mesh.md |
Remove Tailscale references post-migration |
/opt/homelab-context/infrastructure/backups.md |
New strategy |
/opt/homelab-context/infrastructure/monitoring.md |
Enable missing checks |
/opt/homelab-context/containers/103-paperless.md |
Rootfs resize |
/opt/homelab-context/containers/118-elementsynapse.md |
Rootfs resize |
/opt/homelab-context/containers/105-apps.md |
Resource bump, logging addition |
/opt/homelab-context/containers/index.md |
Updated resource table |
.sops.yaml |
New age pubkeys for enrolled LXCs |
Verification
Each phase ends with a verification milestone:
- Phase 1:
claudio-monitortriggers on paperless disk → confirmed alert. Backup of gitea data lands in B2 (or equivalent). DNS query from any LXC returns Technitium answer. - Phase 2:
netbird statusshows all LXCs connected. Tailscale not running on PVE host.curl auth.hubris.networkfrom any LXC resolves correctly without /etc/hosts. - Phase 3: Every LXC has
/opt/homelab-context/+/etc/age/key.txt. MCP tools return valid host info for all enrolled LXC names.journalctlshows promtail shipping to Loki. - Phase 4:
claudio-monitorshows disk growth trend. apps (105) can run all 6+ services without OOM.
Risks & tradeoffs
- Netbird migration window: All LXCs will briefly lose mesh connectivity during the Tailscale→Netbird cutover. Schedule in off-hours.
- Backup cost: B2/e2 costs ~$5/month for ~500 GB. The USB drive was free but unstable — trade money for reliability.
- DNS consolidation: Removing the NetBird managed DNS zone means any NetBird-specific names stop resolving for hubris.network — verify nothing depends on that path.
- Loki on apps (105): Adds another container to an already-loaded host. May need to bump resources before deploying.
- Agent enrollment on every LXC: Each enrollment creates an age keypair
and commits a pubkey to inventory. Process is scriptable via
homelab client addbut still takes ~2 min per host for verification.
Open questions
- Is the USB backup drive still physically attached to hubris? If not, the simplest "re-enable" path requires physically re-attaching it.
- Authentik is now on the VPS — is LXC 124 (old Authentik) still running or was it fully decommissioned? The dns.md changelog says "shut down" but index.md lists it as "running".
- What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure — needed to plan rootfs resizes safely.
- Does the user want to keep Tailscale on any host for a specific reason, or is full Netbird migration the clear goal?