- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
271 lines
12 KiB
Markdown
271 lines
12 KiB
Markdown
# Homelab structure revision & improvement plan
|
|
|
|
## Goal
|
|
|
|
Identify structural issues in the current hubris homelab topology and propose an
|
|
actionable improvement roadmap — DNS consolidation, monitoring gaps, backup
|
|
recovery, mesh completion, resource rightsizing, and operational hygiene.
|
|
|
|
---
|
|
|
|
## Current state summary
|
|
|
|
| Dimension | Status |
|
|
|-----------|--------|
|
|
| Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) |
|
|
| LXCs | 14 active, 1 retired (124), 1 new (107 dns) |
|
|
| VMs | HAOS (108), ZimaOS (100) |
|
|
| Workstations | mac-mini (macOS), republic-laptop, ludo-mini |
|
|
| VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn |
|
|
| Switch | SODOLA 5-Port 2.5Gbit (L2, flat) |
|
|
| Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing |
|
|
| Mesh | Netbird + Tailscale (migrating), ~mixed state |
|
|
| DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS |
|
|
| Backups | DISABLED since 2026-04-22 |
|
|
| Monitoring | claudio-monitor (host health only, no apt/docker checks yet) |
|
|
| Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) |
|
|
|
|
---
|
|
|
|
## Issues identified
|
|
|
|
### 1. Three overlapping DNS sources (highest risk)
|
|
|
|
**Problem:** Technitium on CT 107, NetBird managed DNS zone, and the public
|
|
IONOS wildcard all answer `*.hubris.network` queries. The dns.md doc still
|
|
references the old dnsmasq on LXC 124 (though the change log says it moved).
|
|
NetBird's managed DNS bypasses Technitium entirely for some app names — there
|
|
is no single source of truth for DNS.
|
|
|
|
**Risk:** Mismatched answers → services unreachable → "works on some clients
|
|
but not others" debugging sessions. Already cost time when `auth.hubris.network`
|
|
re-pointed to the VPS.
|
|
|
|
**Proposal:**
|
|
- Phase out NetBird managed DNS zone for `hubris.network` — Technitium is the
|
|
authoritative answerer for mesh & LAN clients
|
|
- Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale
|
|
MagicDNS fallbacks)
|
|
- Document the full authoritative chain: Technitium → upstream forwarders → public
|
|
- Track Technitium config in git (dtoro/technitium-config or equivalent)
|
|
|
|
### 2. Backups disabled with no alternative (data loss risk)
|
|
|
|
**Problem:** The only backup was restic to an external USB that caused host
|
|
crashes. It was disabled 2026-04-22 as an A/B test — host stability was
|
|
confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The
|
|
drive is still removed.
|
|
|
|
**Risk:** `/mnt/library` (~429 GB of irreplaceable data: photos, documents,
|
|
git repos, gitea data) has no off-host copy. Single-disk failure = total loss.
|
|
|
|
**Proposal:**
|
|
- Re-evaluate the USB drive stability with the new SODOLA switch topology
|
|
(direct rear USB 3.0 port, no hub chain)
|
|
- OR adopt a cloud-backed strategy: `rclone`-to-Hetzner Storage Box or
|
|
Backblaze B2 for the irreplaceable subset (docs, photos, gitea data)
|
|
- OR use hubris's own `zfs send` to a second host/disk if ZFS is feasible
|
|
- Minimum viable: at minimum restore gitea backups + sops-encrypted secrets
|
|
via an off-site cron (cheap B2 bucket)
|
|
|
|
### 3. Mesh migration still incomplete
|
|
|
|
**Problem:** Most LXCs still use Tailscale. This forces per-LXC DNS workarounds
|
|
(/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel.
|
|
Mesh migration doc (mesh.md) is comprehensive but execution stalled.
|
|
|
|
**Proposal:**
|
|
- Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
|
|
- Remove Tailscale from the PVE host
|
|
- Ensure all LXCs resolve `*.hubris.network` via Technitium (no more /etc/hosts
|
|
overrides)
|
|
- Document Netbird client on each LXC (netbird version, setup key rotation)
|
|
|
|
### 4. LXC resource imbalance & disk pressure
|
|
|
|
**Problem:**
|
|
| LXC | Cores | RAM | Rootfs | Disk usage |
|
|
|-----|-------|-----|--------|------------|
|
|
| mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable |
|
|
| arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK |
|
|
| nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK |
|
|
| apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight |
|
|
| paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk |
|
|
| elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk |
|
|
| gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate |
|
|
| caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy |
|
|
|
|
Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105)
|
|
is undersized for 6+ services.
|
|
|
|
**Proposal:**
|
|
- Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
|
|
- Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
|
|
- Enable claudio-monitor's disk check to alert before next crisis
|
|
|
|
### 5. VPS is a single point of failure
|
|
|
|
**Problem:** One IONOS VM runs netbird management (control plane), traefik
|
|
(public ingress), authentik (identity), and coturn (TURN relay). If it goes
|
|
down: no remote mesh, no public services, no auth.
|
|
|
|
**Proposal:**
|
|
- Document a VPS recovery runbook (how to restore from a known-working backup)
|
|
- Consider splitting authentik into a separate host or at minimum having a
|
|
standby configuration
|
|
- Not a high priority (the VPS has been stable) but worth documenting the
|
|
blast radius and recovery path
|
|
|
|
### 6. No centralized logging
|
|
|
|
**Problem:** Each LXC has independent journald. Cross-service debugging
|
|
involves hopping between `pct exec <id> -- journalctl -u <service>`. There is
|
|
no aggregation or retention.
|
|
|
|
**Proposal:**
|
|
- Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
|
|
- Ship logs to a central Loki instance on apps (105) or a new small LXC
|
|
- Grafana dashboard optional — even a simple `logcli` query saves time
|
|
|
|
### 7. Agent enrollment incomplete
|
|
|
|
**Problem:** Only hubris, apps, and republic-laptop are enrolled in the
|
|
homelab-context system (age keys, sync timers, MCP access). mac-mini,
|
|
ludo-mini, claudio-bot, and all other LXCs are not.
|
|
|
|
**Proposal:**
|
|
- Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
|
|
- Enroll mac-mini (macOS — exercises the launchd timer path)
|
|
- Enroll ludo-mini (needs SSH user config in inventory first)
|
|
- Wire claudio-bot into inventory-aware queries
|
|
|
|
### 8. Configuration drift on untracked configs
|
|
|
|
**Problem:** Technitium config, dnsmasq (legacy), and several service-specific
|
|
configs are not git-tracked.
|
|
|
|
**Proposal:**
|
|
- Track Technitium zone backup + compose config in a git repo
|
|
- Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy
|
|
|
|
### 9. No capacity planning / resource monitoring
|
|
|
|
**Problem:** No trend data on CPU, RAM, or disk growth. The 86% disk alerts
|
|
were discovered reactively. Rootfs resize is painful (requires Proxmox stop +
|
|
resize + growfs inside).
|
|
|
|
**Proposal:**
|
|
- Enable the missing claudio-monitor checks (disk growth trend, apt upgradable
|
|
counts, docker image drift)
|
|
- Set up a simple Prometheus + node_exporter on hubris or use the PVE API
|
|
directly
|
|
- At minimum, surface disk usage in the existing homelab-mcp management tools
|
|
|
|
### 10. No standard deploy / orchestration for bare-metal LXCs
|
|
|
|
**Problem:** Some services are bare-metal CLI apps (sophia, claudio-bot),
|
|
some are Docker on apps (105), some are Portainer-managed. No consistent
|
|
deploy pattern means every new service reinvents the deployment.
|
|
|
|
**Proposal:**
|
|
- Don't over-engineer this — the current pragmatism works
|
|
- Just document the decision tree:
|
|
- Needs `/mnt/library` mount + heavy I/O → dedicated LXC
|
|
- Small stateless web service → Docker on apps (105)
|
|
- Media stack → dedicated LXC (arriman, jellyfin)
|
|
- Everything else → judge by complexity
|
|
|
|
---
|
|
|
|
## Phased implementation plan
|
|
|
|
### Phase 1 — Critical fixes (this week)
|
|
1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
|
|
2. Enable claudio-monitor disk check + disk-growth alerting
|
|
3. Pick one backup strategy and implement minimum viable (e.g. nightly
|
|
gitea dump + sops-encrypted secrets to B2 via rclone)
|
|
4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone
|
|
for hubris.network)
|
|
|
|
### Phase 2 — Mesh consolidation (next week)
|
|
5. Batch-migrate remaining LXCs from Tailscale to Netbird
|
|
6. Remove Tailscale from PVE host
|
|
7. Remove all per-LXC /etc/hosts DNS overrides
|
|
8. Update DNS documentation to reflect Technitium as single source
|
|
|
|
### Phase 3 — Agent enrollment & logging (next 2 weeks)
|
|
9. Enroll all LXCs in homelab-context (age keys, sync timers)
|
|
10. Enroll mac-mini (macOS launchd path — exercises untested code path)
|
|
11. Enroll ludo-mini
|
|
12. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs
|
|
|
|
### Phase 4 — Resource & monitoring hardening (next month)
|
|
13. Resize apps (105) rootfs, bump RAM
|
|
14. Deploy Prometheus + node_exporter or equivalent for trend data
|
|
15. Track Technitium config in git with auto-deploy
|
|
16. Write VPS recovery runbook
|
|
|
|
### Phase 5 — Drive re-evaluation (optional, behind host-stability gate)
|
|
17. Re-attach USB backup drive with the new SODOLA topology (direct port)
|
|
18. If stable for 7 days, re-enable restic backup schedule (chunked)
|
|
19. If not stable, finalize cloud backup as permanent strategy
|
|
|
|
---
|
|
|
|
## Files likely to change
|
|
|
|
| Path | Change |
|
|
|------|--------|
|
|
| `/opt/homelab-context/inventory.yaml` | LXCs enrolled, resource updates |
|
|
| `/opt/homelab-context/secrets/*.yaml` | New recipients for enrolled LXCs |
|
|
| `/opt/homelab-context/infrastructure/dns.md` | Reflect Technitium as sole source |
|
|
| `/opt/homelab-context/infrastructure/mesh.md` | Remove Tailscale references post-migration |
|
|
| `/opt/homelab-context/infrastructure/backups.md` | New strategy |
|
|
| `/opt/homelab-context/infrastructure/monitoring.md` | Enable missing checks |
|
|
| `/opt/homelab-context/containers/103-paperless.md` | Rootfs resize |
|
|
| `/opt/homelab-context/containers/118-elementsynapse.md` | Rootfs resize |
|
|
| `/opt/homelab-context/containers/105-apps.md` | Resource bump, logging addition |
|
|
| `/opt/homelab-context/containers/index.md` | Updated resource table |
|
|
| `.sops.yaml` | New age pubkeys for enrolled LXCs |
|
|
|
|
## Verification
|
|
|
|
Each phase ends with a verification milestone:
|
|
- Phase 1: `claudio-monitor` triggers on paperless disk → confirmed alert. Backup
|
|
of gitea data lands in B2 (or equivalent). DNS query from any LXC returns
|
|
Technitium answer.
|
|
- Phase 2: `netbird status` shows all LXCs connected. Tailscale not running on
|
|
PVE host. `curl auth.hubris.network` from any LXC resolves correctly without
|
|
/etc/hosts.
|
|
- Phase 3: Every LXC has `/opt/homelab-context/` + `/etc/age/key.txt`. MCP
|
|
tools return valid host info for all enrolled LXC names. `journalctl` shows
|
|
promtail shipping to Loki.
|
|
- Phase 4: `claudio-monitor` shows disk growth trend. apps (105) can run all
|
|
6+ services without OOM.
|
|
|
|
## Risks & tradeoffs
|
|
|
|
- **Netbird migration window:** All LXCs will briefly lose mesh connectivity
|
|
during the Tailscale→Netbird cutover. Schedule in off-hours.
|
|
- **Backup cost:** B2/e2 costs ~$5/month for ~500 GB. The USB drive was free
|
|
but unstable — trade money for reliability.
|
|
- **DNS consolidation:** Removing the NetBird managed DNS zone means any
|
|
NetBird-specific names stop resolving for hubris.network — verify nothing
|
|
depends on that path.
|
|
- **Loki on apps (105):** Adds another container to an already-loaded host.
|
|
May need to bump resources before deploying.
|
|
- **Agent enrollment on every LXC:** Each enrollment creates an age keypair
|
|
and commits a pubkey to inventory. Process is scriptable via `homelab client
|
|
add` but still takes ~2 min per host for verification.
|
|
|
|
## Open questions
|
|
|
|
1. Is the USB backup drive still physically attached to hubris? If not, the
|
|
simplest "re-enable" path requires physically re-attaching it.
|
|
2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or
|
|
was it fully decommissioned? The dns.md changelog says "shut down" but
|
|
index.md lists it as "running".
|
|
3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure
|
|
— needed to plan rootfs resizes safely.
|
|
4. Does the user want to keep Tailscale on any host for a specific reason, or
|
|
is full Netbird migration the clear goal? |