investigations: add moonlight/sunshine WiFi jitter report + index entry
- New investigation doc: 2026-06-03-moonlight-sunshine-wifi-jitter.md - Updated investigations/index.md with link and status - Added .gitignore for .DS_Store - Saved Hermes planning docs from recent sessions
This commit is contained in:
271
.hermes/plans/2026-06-03_150000-homelab-structure-revision.md
Normal file
271
.hermes/plans/2026-06-03_150000-homelab-structure-revision.md
Normal file
@@ -0,0 +1,271 @@
|
||||
# Homelab structure revision & improvement plan
|
||||
|
||||
## Goal
|
||||
|
||||
Identify structural issues in the current hubris homelab topology and propose an
|
||||
actionable improvement roadmap — DNS consolidation, monitoring gaps, backup
|
||||
recovery, mesh completion, resource rightsizing, and operational hygiene.
|
||||
|
||||
---
|
||||
|
||||
## Current state summary
|
||||
|
||||
| Dimension | Status |
|
||||
|-----------|--------|
|
||||
| Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) |
|
||||
| LXCs | 14 active, 1 retired (124), 1 new (107 dns) |
|
||||
| VMs | HAOS (108), ZimaOS (100) |
|
||||
| Workstations | mac-mini (macOS), republic-laptop, ludo-mini |
|
||||
| VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn |
|
||||
| Switch | SODOLA 5-Port 2.5Gbit (L2, flat) |
|
||||
| Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing |
|
||||
| Mesh | Netbird + Tailscale (migrating), ~mixed state |
|
||||
| DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS |
|
||||
| Backups | DISABLED since 2026-04-22 |
|
||||
| Monitoring | claudio-monitor (host health only, no apt/docker checks yet) |
|
||||
| Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) |
|
||||
|
||||
---
|
||||
|
||||
## Issues identified
|
||||
|
||||
### 1. Three overlapping DNS sources (highest risk)
|
||||
|
||||
**Problem:** Technitium on CT 107, NetBird managed DNS zone, and the public
|
||||
IONOS wildcard all answer `*.hubris.network` queries. The dns.md doc still
|
||||
references the old dnsmasq on LXC 124 (though the change log says it moved).
|
||||
NetBird's managed DNS bypasses Technitium entirely for some app names — there
|
||||
is no single source of truth for DNS.
|
||||
|
||||
**Risk:** Mismatched answers → services unreachable → "works on some clients
|
||||
but not others" debugging sessions. Already cost time when `auth.hubris.network`
|
||||
re-pointed to the VPS.
|
||||
|
||||
**Proposal:**
|
||||
- Phase out NetBird managed DNS zone for `hubris.network` — Technitium is the
|
||||
authoritative answerer for mesh & LAN clients
|
||||
- Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale
|
||||
MagicDNS fallbacks)
|
||||
- Document the full authoritative chain: Technitium → upstream forwarders → public
|
||||
- Track Technitium config in git (dtoro/technitium-config or equivalent)
|
||||
|
||||
### 2. Backups disabled with no alternative (data loss risk)
|
||||
|
||||
**Problem:** The only backup was restic to an external USB that caused host
|
||||
crashes. It was disabled 2026-04-22 as an A/B test — host stability was
|
||||
confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The
|
||||
drive is still removed.
|
||||
|
||||
**Risk:** `/mnt/library` (~429 GB of irreplaceable data: photos, documents,
|
||||
git repos, gitea data) has no off-host copy. Single-disk failure = total loss.
|
||||
|
||||
**Proposal:**
|
||||
- Re-evaluate the USB drive stability with the new SODOLA switch topology
|
||||
(direct rear USB 3.0 port, no hub chain)
|
||||
- OR adopt a cloud-backed strategy: `rclone`-to-Hetzner Storage Box or
|
||||
Backblaze B2 for the irreplaceable subset (docs, photos, gitea data)
|
||||
- OR use hubris's own `zfs send` to a second host/disk if ZFS is feasible
|
||||
- Minimum viable: at minimum restore gitea backups + sops-encrypted secrets
|
||||
via an off-site cron (cheap B2 bucket)
|
||||
|
||||
### 3. Mesh migration still incomplete
|
||||
|
||||
**Problem:** Most LXCs still use Tailscale. This forces per-LXC DNS workarounds
|
||||
(/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel.
|
||||
Mesh migration doc (mesh.md) is comprehensive but execution stalled.
|
||||
|
||||
**Proposal:**
|
||||
- Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
|
||||
- Remove Tailscale from the PVE host
|
||||
- Ensure all LXCs resolve `*.hubris.network` via Technitium (no more /etc/hosts
|
||||
overrides)
|
||||
- Document Netbird client on each LXC (netbird version, setup key rotation)
|
||||
|
||||
### 4. LXC resource imbalance & disk pressure
|
||||
|
||||
**Problem:**
|
||||
| LXC | Cores | RAM | Rootfs | Disk usage |
|
||||
|-----|-------|-----|--------|------------|
|
||||
| mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable |
|
||||
| arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK |
|
||||
| nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK |
|
||||
| apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight |
|
||||
| paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk |
|
||||
| elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk |
|
||||
| gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate |
|
||||
| caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy |
|
||||
|
||||
Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105)
|
||||
is undersized for 6+ services.
|
||||
|
||||
**Proposal:**
|
||||
- Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
|
||||
- Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
|
||||
- Enable claudio-monitor's disk check to alert before next crisis
|
||||
|
||||
### 5. VPS is a single point of failure
|
||||
|
||||
**Problem:** One IONOS VM runs netbird management (control plane), traefik
|
||||
(public ingress), authentik (identity), and coturn (TURN relay). If it goes
|
||||
down: no remote mesh, no public services, no auth.
|
||||
|
||||
**Proposal:**
|
||||
- Document a VPS recovery runbook (how to restore from a known-working backup)
|
||||
- Consider splitting authentik into a separate host or at minimum having a
|
||||
standby configuration
|
||||
- Not a high priority (the VPS has been stable) but worth documenting the
|
||||
blast radius and recovery path
|
||||
|
||||
### 6. No centralized logging
|
||||
|
||||
**Problem:** Each LXC has independent journald. Cross-service debugging
|
||||
involves hopping between `pct exec <id> -- journalctl -u <service>`. There is
|
||||
no aggregation or retention.
|
||||
|
||||
**Proposal:**
|
||||
- Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
|
||||
- Ship logs to a central Loki instance on apps (105) or a new small LXC
|
||||
- Grafana dashboard optional — even a simple `logcli` query saves time
|
||||
|
||||
### 7. Agent enrollment incomplete
|
||||
|
||||
**Problem:** Only hubris, apps, and republic-laptop are enrolled in the
|
||||
homelab-context system (age keys, sync timers, MCP access). mac-mini,
|
||||
ludo-mini, claudio-bot, and all other LXCs are not.
|
||||
|
||||
**Proposal:**
|
||||
- Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
|
||||
- Enroll mac-mini (macOS — exercises the launchd timer path)
|
||||
- Enroll ludo-mini (needs SSH user config in inventory first)
|
||||
- Wire claudio-bot into inventory-aware queries
|
||||
|
||||
### 8. Configuration drift on untracked configs
|
||||
|
||||
**Problem:** Technitium config, dnsmasq (legacy), and several service-specific
|
||||
configs are not git-tracked.
|
||||
|
||||
**Proposal:**
|
||||
- Track Technitium zone backup + compose config in a git repo
|
||||
- Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy
|
||||
|
||||
### 9. No capacity planning / resource monitoring
|
||||
|
||||
**Problem:** No trend data on CPU, RAM, or disk growth. The 86% disk alerts
|
||||
were discovered reactively. Rootfs resize is painful (requires Proxmox stop +
|
||||
resize + growfs inside).
|
||||
|
||||
**Proposal:**
|
||||
- Enable the missing claudio-monitor checks (disk growth trend, apt upgradable
|
||||
counts, docker image drift)
|
||||
- Set up a simple Prometheus + node_exporter on hubris or use the PVE API
|
||||
directly
|
||||
- At minimum, surface disk usage in the existing homelab-mcp management tools
|
||||
|
||||
### 10. No standard deploy / orchestration for bare-metal LXCs
|
||||
|
||||
**Problem:** Some services are bare-metal CLI apps (sophia, claudio-bot),
|
||||
some are Docker on apps (105), some are Portainer-managed. No consistent
|
||||
deploy pattern means every new service reinvents the deployment.
|
||||
|
||||
**Proposal:**
|
||||
- Don't over-engineer this — the current pragmatism works
|
||||
- Just document the decision tree:
|
||||
- Needs `/mnt/library` mount + heavy I/O → dedicated LXC
|
||||
- Small stateless web service → Docker on apps (105)
|
||||
- Media stack → dedicated LXC (arriman, jellyfin)
|
||||
- Everything else → judge by complexity
|
||||
|
||||
---
|
||||
|
||||
## Phased implementation plan
|
||||
|
||||
### Phase 1 — Critical fixes (this week)
|
||||
1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
|
||||
2. Enable claudio-monitor disk check + disk-growth alerting
|
||||
3. Pick one backup strategy and implement minimum viable (e.g. nightly
|
||||
gitea dump + sops-encrypted secrets to B2 via rclone)
|
||||
4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone
|
||||
for hubris.network)
|
||||
|
||||
### Phase 2 — Mesh consolidation (next week)
|
||||
5. Batch-migrate remaining LXCs from Tailscale to Netbird
|
||||
6. Remove Tailscale from PVE host
|
||||
7. Remove all per-LXC /etc/hosts DNS overrides
|
||||
8. Update DNS documentation to reflect Technitium as single source
|
||||
|
||||
### Phase 3 — Agent enrollment & logging (next 2 weeks)
|
||||
9. Enroll all LXCs in homelab-context (age keys, sync timers)
|
||||
10. Enroll mac-mini (macOS launchd path — exercises untested code path)
|
||||
11. Enroll ludo-mini
|
||||
12. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs
|
||||
|
||||
### Phase 4 — Resource & monitoring hardening (next month)
|
||||
13. Resize apps (105) rootfs, bump RAM
|
||||
14. Deploy Prometheus + node_exporter or equivalent for trend data
|
||||
15. Track Technitium config in git with auto-deploy
|
||||
16. Write VPS recovery runbook
|
||||
|
||||
### Phase 5 — Drive re-evaluation (optional, behind host-stability gate)
|
||||
17. Re-attach USB backup drive with the new SODOLA topology (direct port)
|
||||
18. If stable for 7 days, re-enable restic backup schedule (chunked)
|
||||
19. If not stable, finalize cloud backup as permanent strategy
|
||||
|
||||
---
|
||||
|
||||
## Files likely to change
|
||||
|
||||
| Path | Change |
|
||||
|------|--------|
|
||||
| `/opt/homelab-context/inventory.yaml` | LXCs enrolled, resource updates |
|
||||
| `/opt/homelab-context/secrets/*.yaml` | New recipients for enrolled LXCs |
|
||||
| `/opt/homelab-context/infrastructure/dns.md` | Reflect Technitium as sole source |
|
||||
| `/opt/homelab-context/infrastructure/mesh.md` | Remove Tailscale references post-migration |
|
||||
| `/opt/homelab-context/infrastructure/backups.md` | New strategy |
|
||||
| `/opt/homelab-context/infrastructure/monitoring.md` | Enable missing checks |
|
||||
| `/opt/homelab-context/containers/103-paperless.md` | Rootfs resize |
|
||||
| `/opt/homelab-context/containers/118-elementsynapse.md` | Rootfs resize |
|
||||
| `/opt/homelab-context/containers/105-apps.md` | Resource bump, logging addition |
|
||||
| `/opt/homelab-context/containers/index.md` | Updated resource table |
|
||||
| `.sops.yaml` | New age pubkeys for enrolled LXCs |
|
||||
|
||||
## Verification
|
||||
|
||||
Each phase ends with a verification milestone:
|
||||
- Phase 1: `claudio-monitor` triggers on paperless disk → confirmed alert. Backup
|
||||
of gitea data lands in B2 (or equivalent). DNS query from any LXC returns
|
||||
Technitium answer.
|
||||
- Phase 2: `netbird status` shows all LXCs connected. Tailscale not running on
|
||||
PVE host. `curl auth.hubris.network` from any LXC resolves correctly without
|
||||
/etc/hosts.
|
||||
- Phase 3: Every LXC has `/opt/homelab-context/` + `/etc/age/key.txt`. MCP
|
||||
tools return valid host info for all enrolled LXC names. `journalctl` shows
|
||||
promtail shipping to Loki.
|
||||
- Phase 4: `claudio-monitor` shows disk growth trend. apps (105) can run all
|
||||
6+ services without OOM.
|
||||
|
||||
## Risks & tradeoffs
|
||||
|
||||
- **Netbird migration window:** All LXCs will briefly lose mesh connectivity
|
||||
during the Tailscale→Netbird cutover. Schedule in off-hours.
|
||||
- **Backup cost:** B2/e2 costs ~$5/month for ~500 GB. The USB drive was free
|
||||
but unstable — trade money for reliability.
|
||||
- **DNS consolidation:** Removing the NetBird managed DNS zone means any
|
||||
NetBird-specific names stop resolving for hubris.network — verify nothing
|
||||
depends on that path.
|
||||
- **Loki on apps (105):** Adds another container to an already-loaded host.
|
||||
May need to bump resources before deploying.
|
||||
- **Agent enrollment on every LXC:** Each enrollment creates an age keypair
|
||||
and commits a pubkey to inventory. Process is scriptable via `homelab client
|
||||
add` but still takes ~2 min per host for verification.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is the USB backup drive still physically attached to hubris? If not, the
|
||||
simplest "re-enable" path requires physically re-attaching it.
|
||||
2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or
|
||||
was it fully decommissioned? The dns.md changelog says "shut down" but
|
||||
index.md lists it as "running".
|
||||
3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure
|
||||
— needed to plan rootfs resizes safely.
|
||||
4. Does the user want to keep Tailscale on any host for a specific reason, or
|
||||
is full Netbird migration the clear goal?
|
||||
Reference in New Issue
Block a user