- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
129 lines
6.4 KiB
Markdown
129 lines
6.4 KiB
Markdown
# Plan: Prevent DHCP IP drift from breaking Caddy backends
|
||
|
||
**Date:** 2026-06-05
|
||
**Slug:** prevent-dhcp-ip-drift
|
||
|
||
---
|
||
|
||
## Goal
|
||
|
||
Eliminate the root cause of services becoming unreachable when DHCP lease renewals change backend IPs that Caddy's `reverse_proxy` directives hardcode.
|
||
|
||
**Triggering incident:** Paperless (LXC 103) and HAOS (VM 108) had DHCP-assigned IPs change from `.130→.243` and `.101→.241` respectively. Caddyfile still pointed at the old IPs → services unreachable from iPhone on Netbird.
|
||
|
||
## Current context
|
||
|
||
### DHCP vs static IP inventory
|
||
|
||
| Machine | Type | PVE ID | Current IP | Allocation method | Status |
|
||
|---------|------|--------|-----------|-------------------|--------|
|
||
| **jellyfin** | LXC | 101 | 192.168.8.206 | Static (`ip=.../24`) | ✅ |
|
||
| **paperless** | LXC | 103 | 192.168.8.243 | **DHCP** (`ip=dhcp`) | ❌ broken, hotfixed to .243 |
|
||
| **gitea** | LXC | 104 | 192.168.8.121 | Static | ✅ |
|
||
| **apps** | LXC | 105 | 192.168.8.205 | Static | ✅ |
|
||
| **nextcloud** | LXC | 114 | 192.168.8.224 | Static | ✅ |
|
||
| **elementsynapse** | LXC | 118 | 192.168.8.239 | Static | ✅ |
|
||
| **mule-images** | LXC | 120 | 192.168.8.136 | Static | ✅ |
|
||
| **caddy** | LXC | 121 | 192.168.8.175 | Static | ✅ |
|
||
| **arriman** | LXC | 122 | 192.168.8.132 | Static | ✅ |
|
||
| **sophia** | LXC | 119 | 192.168.8.157 | Static? | ? (not in 2026-06-02 list) |
|
||
| **nfs-export** | LXC | 102 | 192.168.8.200 | Static? | ? |
|
||
| **plato** | LXC | 126 | 192.168.8.190 | Static? | ? |
|
||
| **HAOS** | VM | 108 | 192.168.8.241 | **DHCP** (VM — OS-managed) | ❌ broken, hotfixed to .241 |
|
||
| **zimaos** | VM | 100 | 192.168.8.195 | DHCP (known stale lease, see note) | ⚠️ open issue |
|
||
| **authentik** | — | — | — | migrated to VPS (external) | N/A |
|
||
|
||
### Infrastructure facts
|
||
|
||
- **DHCP server:** Technitium on CT 107 (192.168.8.2), pool `.241–.254`
|
||
- **Static IP range:** `.101–.239`
|
||
- **DNS:** Split-horizon on Technitium — `*.hubris.network → 192.168.8.175` (Caddy itself)
|
||
- **Caddyfile:** Has 29 `reverse_proxy` directives, all using **hardcoded IP:port**
|
||
- **Caddy reload:** Auto-deployed via webhook on git push to `dtoro/caddy-conf`
|
||
- **Documentation:** `inventory.yaml` updated, `hosts/*.yaml` stale-before-regenerate
|
||
|
||
### Why it happened
|
||
|
||
1. Paperless LXC 103 was **missed** during the 2026-06-02 static-IP migration (Proxmox config still shows `ip=dhcp`)
|
||
2. HAOS VM 108 uses DHCP internally (HAOS manages its own network — can't just `pct set`)
|
||
3. Both got new IPs from the Technitium `.241–.254` pool after a lease renewal
|
||
4. Caddyfile still pointed at the old static-range IPs → connection refused
|
||
|
||
---
|
||
|
||
## Proposed approach
|
||
|
||
Three-layer solution:
|
||
|
||
### Layer 1: Fix the immediate offenders (static assignment)
|
||
|
||
**Paperless LXC 103:**
|
||
- `pct set 103 --net0 name=eth0,bridge=vmbr0,gw=192.168.8.1,hwaddr=BC:24:11:0A:8D:C2,ip=192.168.8.130/24,ip6=auto,type=veth`
|
||
- Inside the LXC, update `/etc/network/interfaces` to match
|
||
- Reboot/restart networking
|
||
|
||
**HAOS VM 108:**
|
||
- Set a **DHCP reservation** in Technitium for the VM's MAC address, pinning it to `192.168.8.101`
|
||
- This avoids needing to reconfigure HAOS internally (which is tedious)
|
||
- Alternatively: use `ha network update` from the HAOS console to set a static IP
|
||
|
||
### Layer 2: Audit and fix all remaining DHCP hosts
|
||
|
||
Check every LXC/VM hosted on hubris:
|
||
```bash
|
||
for ct in $(pct list | awk 'NR>1{print $1}'); do
|
||
echo "=== CT $ct ==="
|
||
pct config $ct | grep "^net0"
|
||
done
|
||
```
|
||
Any with `ip=dhcp` that Caddy reverse-proxies to → convert to static.
|
||
|
||
**Known candidates to check:**
|
||
- CT 102 (nfs-export) — `.200` but not in Caddy. May not need static.
|
||
- CT 119 (sophia) — `.157` — is this static or DHCP? Not sure.
|
||
- CT 126 (plato) — `.190` — same question.
|
||
- VM 100 (zimaos) — `.195` but known to have a stale lease (see 2026-06-03 changelog)
|
||
|
||
### Layer 3: Add validation
|
||
|
||
Create a script that runs periodically (cron or homelab cronjob):
|
||
|
||
**`/opt/homelab-context/scripts/check-caddy-backends.sh`:**
|
||
1. Parse `/etc/caddy/Caddyfile` on CT 121 to extract all `reverse_proxy IP:port` targets
|
||
2. For each `IP:port`, attempt a TCP connect (timeout 3s)
|
||
3. Report any that fail
|
||
|
||
Could also run as a homelab cron job that notifies dtoro on Matrix if a backend is unreachable.
|
||
|
||
This catches any future drift proactively (before a user reports it).
|
||
|
||
### Files likely to change
|
||
|
||
| File | Change |
|
||
|------|--------|
|
||
| `inventory.yaml` | May update paperless/HAOS IPs if we choose different static IPs |
|
||
| LXC 103 Proxmux config (via `pct set`) | Set static IP |
|
||
| Technitium DHCP reservations | Add HAOS reservation |
|
||
| `/etc/caddy/Caddyfile` on CT 121 | Already fixed — only changes again if we re-assign paperless IP to `.130` |
|
||
| `scripts/check-caddy-backends.sh` | New validation script (new file in homelab-context) |
|
||
|
||
### Risks / Tradeoffs
|
||
|
||
- **Pinning paperless to `.130`** — if the LXC was reinstalled since then, `.130` may already be in use. Verify first with `arp-scan` or `nmap`.
|
||
- **HAOS static IP via Technitium reservation** vs **inside HAOS**: Technitium reservation is simpler (no HA config changes), but if HAOS's DHCP lease expires and the Technitium server is down, the reservation won't help. A static IP inside HAOS is more robust but requires poking the HA console.
|
||
- **Validation script false positives** — a service might be legitimately down for maintenance. The script should be a warning, not an alert.
|
||
- **Caddy reload** — each Caddyfile edit triggers an auto-reload via webhook. If the backend is down during reload, Caddy itself stays up (it's just a reverse_proxy target).
|
||
|
||
### Verification
|
||
|
||
1. After setting paperless static: `ssh root@192.168.8.175 "curl -s -o /dev/null -w '%{http_code}' http://192.168.8.130:8000"` → 302
|
||
2. After Technitium HAOS reservation: `curl -s -o /dev/null -w '%{http_code}' http://192.168.8.101:8123` → 200
|
||
3. Run validation script → all targets reachable
|
||
4. Confirm from iPhone: both `paperless.hubris.network` and `home.hubris.network` load
|
||
|
||
### Open questions
|
||
|
||
1. Should paperless go back to `.130` (its original), or stay at `.243` (current)? Going back to `.130` means updating the Caddyfile again, but keeps the static range allocation consistent.
|
||
2. HAOS: Technitium reservation or HAOS-internal static config? Reservation is easier; HAOS-internal is more robust.
|
||
3. Should the Caddyfile validation script run as a homelab cron job, or as a cron on the caddy LXC itself?
|
||
4. ZimaOS (VM 100) — should we also pin its IP while we're at it? |